A managed ingestion tool fits when a vendor API should become one more pipeline: scheduled, monitored and landed in your warehouse beside your other sources. This guide builds that pipeline for a cursor-paginated API in Fivetran and in Airbyte, with the Fokals REST API as the worked example. It covers how the key is sent, how the paging loop ends, which value is the bookmark and what to do when the rate limit answers.
Fokals is delivered direct, by REST API and as bulk files in JSON, JSON Lines or CSV, so the connector below is one you write and run in either tool, and your connector loads the data into your destination. The API reference is the document to build against.
What the API asks of a connector
Every Fokals endpoint is a GET under /api/v1. The behaviours that shape a connector are these.
| Concern | What the API reference says |
|---|---|
| Authentication | A key sent as Authorization: Bearer <key> or in an X-Api-Key header. A key has the scope read, exports or both. |
| Records | Lists answer { "data": [...], "next_cursor": "..." }. |
| Paging | Send next_cursor back as cursor with the same other parameters. The last page has next_cursor set to null. limit runs from 1 to 200, and from 1 to 1000 on exports. |
| Incremental feeds | /changes, /signals and /postings run oldest first from since. /news runs newest first. Where the reference names a default for since, it is seven days ago. |
| Rate limits | Counted per key by minute and by day. Over a limit the answer is 429 with Retry-After in seconds, and a request refused for the daily limit still counts. |
| Bulk | /exports/{product}/{dataset} takes from and to, a format of json, jsonl or csv, and sends the next cursor in the X-Next-Cursor header. |
The change feed is a good first stream. Its rows are the dated events of Technology Changes (company_tech_events) in the marketing stack dataset: a technology added or removed, a market or currency added, a platform replaced, each with the time it was observed.
Choosing the route in each tool
The Fivetran Connector SDK lets you write custom connectors in Python and deploys them as extensions of Fivetran. A connector is a connector.py that declares a Connector object around an update(configuration, state) function and, optionally, a schema(configuration) function, with credentials in configuration.json and dependencies in requirements.txt or pyproject.toml. You run it locally with fivetran debug and deploy it with fivetran deploy. Fivetran then runs the connection on its schedule and hands your code the state it last saved. Its page on function connectors, which run on AWS Lambda, Azure Functions or Google Cloud Functions, says users who signed up on or after 22 July 2025 no longer have access to them, so this guide uses the SDK.
Airbyte's connector development overview lists, among its options, the Connector Builder, the low-code CDK and the Python CDK. The Connector Builder is a no-code tool in the Airbyte interface that generates a YAML manifest and builds source connectors only; the overview recommends it for an API source. The low-code CDK is the same declarative YAML written by hand, with custom Python components where needed. The Python CDK offers the most flexibility and, the overview notes, needs the most code and maintenance. A cursor API with bearer authentication is within what the Builder's forms cover: authentication, streams, pagination, incremental sync and error handling.
Mapping the API onto each tool
The table sets each need beside the setting that meets it. Option names are those on the Airbyte pages and in the Fivetran SDK reference.
| Need | Airbyte Connector Builder | Fivetran Connector SDK |
|---|---|---|
| Send the key | Bearer Token authentication, the key a user input rather than saved in the connector | configuration["api_key"] from configuration.json, sent as a bearer header |
| Find the records | Record Selector, Field Path data | body["data"] |
| Walk the pages | Cursor Pagination: the cursor read from the response body as {{ response['next_cursor'] }} and injected as the request parameter cursor | A loop that sends next_cursor back as cursor until it is null |
| Keep the bookmark | Incremental sync: cursor field observed_at, a start datetime, the start time injected as the request parameter since, and a lookback window | state["since"], saved with op.checkpoint |
| Survive a rate limit | Error handling: the RATE_LIMITED action with the backoff Wait Time Extracted from Response Header, naming Retry-After | Read Retry-After and sleep, or stop the run |
| Name the row | Primary Key id | primary_key in the schema function |
The API reference describes feeds as walkable by keeping the last cursor. The connector here keeps a timestamp instead, because a cursor is tied to the query that made it, and one that does not fit the request is refused with 400 invalid_request. A timestamp works with any limit or filter.
Its cost is an overlap at the boundary, because an incremental read starts at the last value, not after it. A primary key absorbs the overlap when the destination writes by key: Fivetran's op.upsert does, and in Airbyte the sync mode Incremental | Append + Deduped keeps one row per primary key, while Incremental | Append can leave several copies of a record. The incremental sync use case treats bookmarks, idempotent upserts and backfills in more depth.
The Fivetran connector
This connector loads the change feed. The before and after values are objects, so it stores them as JSON text, which is how the bulk files carry objects too: JSON in a single cell.
import json
import time
import requests
from fivetran_connector_sdk import Connector
from fivetran_connector_sdk import Operations as op
TABLE = "company_tech_events"
def schema(configuration: dict):
return [{"table": TABLE, "primary_key": ["id"]}]
def update(configuration: dict, state: dict):
headers = {"Authorization": f"Bearer {configuration['api_key']}"}
since = state.get("since", configuration["start"])
params = {"since": since, "limit": 200}
saved = time.monotonic()
while True:
response = requests.get(
f"{configuration['base_url']}/changes", headers=headers, params=params, timeout=60
)
if response.status_code == 429:
wait = int(response.headers.get("Retry-After", "60"))
if wait > 300:
op.checkpoint(state={"since": since})
return # the next scheduled run resumes from the bookmark
time.sleep(wait)
continue
response.raise_for_status()
body = response.json()
for row in body["data"]:
for key in ("before", "after"):
row[key] = None if row[key] is None else json.dumps(row[key])
op.upsert(table=TABLE, data=row)
since = row["observed_at"]
if time.monotonic() - saved > 600:
op.checkpoint(state={"since": since})
saved = time.monotonic()
if not body["next_cursor"]:
break
params["cursor"] = body["next_cursor"]
op.checkpoint(state={"since": since})
connector = Connector(update=update, schema=schema)In configuration.json, base_url is https://<your Fokals host>/api/v1 and start is the first bookmark. Set start deliberately: the feeds that name a default, such as /signals and /postings, begin seven days back when since is left out, so name the first bookmark you need. The id primary key makes the overlap harmless. Fivetran's data handling page says a defined primary key makes an update change the existing record instead of adding a duplicate, and that without one Fivetran hashes the whole record.
The state management page describes op.checkpoint as telling Fivetran that the data sent so far can be written to the destination, and suggests checkpointing about every ten minutes and not more than once a minute. That is why the loop checkpoints on a timer, not on every page. On a 429 it waits up to five minutes and otherwise checkpoints and ends the run. A refusal for the daily limit still counts, so sleeping and retrying all day only extends the refusal.
The same stream in the Airbyte Builder
Build the stream with these settings, then test it on the last page of a walk.
- Authentication: choose Bearer Token, one of the methods the authentication page lists, and leave the token as a user input. The overview says to pass sensitive credentials as user inputs after publishing, not to hardcode them.
- Stream: set the base URL to
https://<your Fokals host>/api/v1, the path to/changes, the Primary Key toidand the Record Selector's Field Path todata. - Pagination: choose Cursor Pagination, read the cursor from the response body with
{{ response['next_cursor'] }}and inject it as the request parametercursor. The pagination page says the connector stops when no cursor value is found or when a stop condition you write is true, so check the last page, wherenext_cursorisnull. - Incremental sync: set the cursor field to
observed_at, list the datetime format of the API's timestamps (2026-09-22T10:11:12.123456Zhas six fractional digits and aZ), make the start datetime a user input, inject the start time as the parametersinceand add a lookback window. The incremental sync page describes each option. - Error handling: add a response filter for HTTP 429 with the action RATE_LIMITED and the backoff Wait Time Extracted from Response Header, naming
Retry-After, as the error handling page advises for that header. - Publish the connector to your workspace, then enter the key and the start date when you create the source.
- When you connect the source to a destination, choose Incremental | Append + Deduped for the stream, so the primary key keeps one row per change.
Pitfalls that apply to both
Order. /news runs newest first and its cursor walks back toward since, so a run is complete only when it reaches since. In the SDK, checkpoint once, at the end, or take announcements for a period from the Company News export (company_news) instead, where rows are ordered by the dataset's keys and a period can be resumed.
Other feeds. /signals rows carry observed_at like /changes, while /postings and /news rows carry at, so the cursor field changes with the stream. On /postings, postings in the baseline are left out unless baseline=true, because a first observation sets a baseline and is never counted as a change. A posting appears once for each event, opened and closed, so key that stream on id and event.
Shared limits. The limits belong to the key. A connector that shares a key with a notebook or a dashboard draws on the same minute and daily budget, and every response reports what is left in X-RateLimit-Remaining-Day.
Schema changes. The reference says fields are added within a version without notice. Labels and scores are produced under named, frozen versions such as jobs-v2, and a breaking change ships as a new version with at least 90 days' notice. Land label_version in the destination and filter on it.
When a connector is not needed
A connector earns its place when you already run one of these tools and want one place to schedule and watch every source. If the need is a daily or weekly file in a warehouse, the warehouse's own loader may be all you need; API versus bulk files sets out the trade. A connector also lands rows, not a model. Joining them on the Fokals company ID, ISIN or FIGI and testing keys and freshness is the next step, covered in modelling licensed company data with dbt. For the two ingestion tools side by side, see Fivetran versus Airbyte for vendor APIs.
Frequently asked questions
How do I build a custom connector for a REST API in Fivetran?
Use the Connector SDK. Write a Python connector.py that defines update(configuration, state), calls the API, passes each row to op.upsert and saves progress with op.checkpoint. Put credentials in configuration.json, test locally with fivetran debug, then run fivetran deploy. Fivetran then runs the connection on its schedule and gives your code the state it last saved.
Can Airbyte read a cursor-paginated API without code?
Yes, within what the Connector Builder offers. Its forms cover authentication, streams, pagination, incremental sync and error handling, and its Cursor Pagination strategy reads the next-page cursor from the response body or headers and injects it into the next request. The Builder creates source connectors only and writes a YAML manifest that you can publish to your workspace or export.
How does Fokals data reach a Fivetran or Airbyte destination?
Fokals is delivered direct, by REST API and as bulk files in JSON, JSON Lines or CSV. You build a connector against the API in either tool, as in this guide, or load the bulk files into your destination with its own loader.
How should a connector handle an API rate limit?
Read the limit the API reports and wait it out. The Fokals API answers 429 with a Retry-After header in seconds, and a request refused for the daily limit still counts, so retrying at once only extends the refusal. In Airbyte use the RATE_LIMITED action with the backoff that reads the wait from a response header. In the Fivetran SDK read the header yourself, sleep for short waits and end the run, after a checkpoint, for long ones.
Should I ingest a vendor API or load its bulk files?
Use the API for feeds you want to follow, such as the change, signal and posting feeds that run oldest first from a time you set. Use bulk exports for a period or a first load: they take from and to, return JSON, JSON Lines or CSV, and come with a manifest naming sources, period, label versions and licence. The two combine: load the period you need in bulk, then follow the feed from its end.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.
What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.