Glossary

Point-in-time data

Data stored as it stood on each past date, with no later corrections. How the two dates of a record work, an as-of query and the mistakes that break a back-test.

Updated 5 October 20262 min read

Point-in-time data stores each value with the date it became known, so a query for any past date returns what was available then rather than what was learned later. It keeps later restatements, corrections and backfills out of the past, which makes a back-test fair.

Two dates on every record

A record has the period it describes and the moment it became known. For example, a quarter ends on 31 March and its figures are published on 5 May, so a model that joins on the period end sees May's figures in April. Point-in-time storage keeps both dates and answers a query by the second one. Databases call this a bitemporal design.

An as-of query

The query below returns the daily rows for the days closed before 1 October 2026, leaving out any row written late. Acme Robotics, an illustrative company, stands for the company id.

select day, open_postings, new_postings, closed_postings
from company_hiring_daily
where company_id = :acme_id
  and day < date '2026-10-01'
  and reconstructed = false;

Common mistakes

  • Joining on the period a value describes, not the date it was published.
  • Correcting a table in place, which erases what was known before the correction.
  • Keeping a label's name when its definition changes. A new definition needs a new label version.
  • Using today's list of companies for a past date, which is survivorship bias.

In Fokals data

Every observation carries its source and the time it was observed. Daily and weekly tables are written once, after the period closes, and never revised, so Fokals data is point-in-time by construction. A flag marks a period written more than seven days after it closed, so a test can leave it out. The first observation of a website or job board sets a baseline and is never counted as a change.

Labels and scores are produced under frozen version names, and a breaking change ships as a new version with at least 90 days' notice, as the methodology sets out. Why we write each table once explains the design, and the use case for back-tests shows how to store and query it.

Frequently asked questions

What is the difference between point-in-time data and historical data?

Historical data is any record of the past. Point-in-time data is historical data stored as it stood on each past date. A series that has been restated or backfilled shows the best present knowledge of the past, which nobody had then. A point-in-time series shows what was known at each date, which is what a back-test needs.

Why does point-in-time data matter for back-testing?

A back-test asks what a rule would have done. If its inputs include values that were corrected or added later, the rule acts on knowledge it could not have had, and the result flatters it. That error is look-ahead bias. Point-in-time inputs remove it at the source, instead of relying on the analyst to lag every field by hand.

How can you tell whether a dataset is point-in-time?

Ask whether a period's rows are ever changed after they are first written, whether each record carries the time it was observed, and whether definitions change under the same name. Then compare: take the same past period from two downloads some weeks apart. If the values differ and nothing marks the later ones, the dataset is not point-in-time.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.