Incremental sync keeps a local copy of a dataset current by fetching only the records that are new or changed since the last run, instead of reloading everything. A bookmark, such as a timestamp or a cursor, records how far the previous run got, so each run starts where the last one ended.
A daily run, step by step
- Read the saved bookmark. If there is none, start from the date you choose for the first load.
- Request the feed from the bookmark, oldest record first, page by page.
- Write each page to a staging table, then upsert it into the target on the record key.
- Save the last cursor as the new bookmark only after the write has committed.
- Log the run: the bookmark before and after, and the rows received and written.
The write must be repeatable, so that a retried page or an overlapping run changes nothing. For a table written once per period, that is an insert that ignores keys already present:
insert into company_hiring_daily
select * from staging_company_hiring_daily
on conflict (company_id, day) do nothing;This is PostgreSQL syntax. It assumes a staging table with the same columns and a unique constraint on the company ID and the day.
Where it goes wrong
- A bookmark taken from your own clock instead of the source's cursor leaves gaps or overlaps when the clocks differ.
- A bookmark saved before the data is committed loses a page when the job fails.
- A source that revises old records needs an update time to find them, and a feed that only appends will not carry the changes.
- The first load and the first increment need an overlap, or the rows between them are never fetched.
In Fokals data
Feeds run oldest first from a time you set, so the last cursor is the bookmark. Daily and weekly tables are written once, after the period closes, and are not revised, so for them an incremental sync only appends, and a repeated run adds nothing when you key on the record. Every record carries the time it was observed, and a reconstructed flag marks a period written more than seven days after it closed, so keep that field to tell a late period from an on-time one.
For a first load, take history from a bulk export or from the feed with an early start time, then continue from where it ended. For the parameters, see the API documentation. The use case on keeping a warehouse in step builds the full loop.
Related terms
- Cursor pagination: how a feed returns records in pages and marks its position.
- Bulk export: the usual route for the first load.
- Rate limit: the cap on requests that a long backfill must respect.
- Data freshness: how soon a change at the source reaches your copy.
Frequently asked questions
What is the difference between incremental sync and a full refresh?
A full refresh reloads the whole dataset on every run, so it always matches the source but costs time and request limits in proportion to the dataset's size. An incremental sync fetches only what is new since the bookmark, so each run is small. Its risk is drift: a missed page or a wrong bookmark leaves a gap. Teams therefore reconcile counts regularly, or run an occasional full refresh.
How do I do the first load before starting an incremental sync?
Load the history from a bulk export, or from the feed with an early start time, and note where it ended. Start the incremental feed from that point with a small overlap. A load that ignores rows it already holds will absorb the duplicates the overlap brings. Then compare the row counts for the overlapping days with the source before you trust the bookmark.
How do I stop incremental sync from creating duplicates?
Key every table on the columns that identify a record, and write with an upsert or an insert that ignores existing keys. A retried page, an overlap window or a repeated run then adds nothing. Save the bookmark only after the write commits, so a failure repeats work and never skips it.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.