A point-in-time dataset says what was known on each date and never rewrites it, so a back-test reads the past as an investor would have read it then. This guide is for a quantitative researcher who wants to test company-level signals such as hiring, technology changes and intent scores. It shows what makes the Fokals record point-in-time, how to give each dataset a knowledge date, how to store the loads, how to query as of a date and how to build a list free of survivors.
Three ways company data leaks the future
Company datasets fail a back-test in three ways, and each needs its own defence.
- Look-ahead. The test reads a value that was corrected, restated or backfilled after the date it describes, or a value dated by when the event happened rather than when anyone could see it. Look-ahead bias of this kind inflates a signal that nobody could have traded.
- Survivorship. The list is built from the companies that exist today, so those that were delisted or dropped during the test are missing, and with them the cases where the signal was wrong. This is survivorship bias.
- Baseline artefacts. The first time a site or a job board is observed, everything on it is new to the dataset and none of it is new to the world. A series built from first-seen dates jumps whenever a company enters the data.
The properties of the record
Five properties of the Fokals record do most of the work against those three failures.
- Every observation is dated. Each record carries the time it was observed, so you can tell when a fact entered the data as well as when it happened.
- Written once. Daily and weekly datasets are written after the period closes, so a weekly row read next year is exactly the row written when the week closed.
- Never revised. A row is not corrected, restated or backfilled later. A back-test run today and one run in a year read the same values.
- A first observation is a baseline. The first observation of a site or job board is stored as a baseline and never counted as a change, so a company entering the data does not look like a burst of activity.
- Versioned definitions. Labels and scores are produced under named, frozen versions, and a breaking change ships as a new version with at least 90 days' notice.
One flag marks any row that was written later than the normal schedule. reconstructed=true sits on a daily or weekly row whose period was written more than seven days after it closed. The row is still written once and never changed, but it was not available when the period ended, so a strict test leaves it out and reports how many rows that removed.
A knowledge date for every dataset
A row's own date is not always the day it could be known. Decide the knowledge date for each dataset once, write it down, and add a lag you can defend. The table gives the earliest date before your lag.
| Dataset | Date column | Earliest known | Reading note |
|---|---|---|---|
| Hiring Activity | day | The next UTC day | Clean once reconstructed rows are out |
| Sales Team Metrics, Intent Scores | week_start | Seven days after week_start | The Monday opens the week, it does not close it |
| Market Series | as_of | The day after as_of | The windows end on that Sunday |
| Job Postings | first_seen_at | first_seen_at | posted_at is the board's date and can be earlier; found_on_first_read marks postings already open at the baseline |
| Technology Changes | observed_at | observed_at | A site's baseline observation writes no events |
| Company Signals | observed_at | The date plus a lag | Carries the posting's own date or the filing date where there is one |
| Company News | at | The date plus a lag | The company's publication date; announcements refresh daily to every three days |
| Company Funding | filed_at | filed_at | first_sale is earlier, normally by no more than 15 days |
Weekly Intent Scores need one more rule. A score is computed from the dated signals of the trailing 90 days and written once, so the stored row is the evidence-backed score as of the week's close, with the dated signals behind it. Read a past week from the stored row. A rebuild from Company Signals today would pick up signals observed after the week closed.
Storing it
Keep the data in three layers.
- Raw loads. Append every file or page you pull, untouched, with two columns of your own:
loaded_at, the time your pipeline received the row, and the file or cursor it came from. Daily and weekly rows are written once, so a period loaded twice gives the same rows, and a unique key on each natural key makes the load safe to repeat:company_idwithdayfor Hiring Activity, andcompany_idwithweek_startfor the weekly datasets, plustopicfor Intent Scores. - A knowledge calendar. One view per dataset that adds
known_at, built from the table above plus your lag. - As-of views. The only layer research code reads. Given a decision date, it returns rows whose
known_atis on or before it.
Key your own tables on company_id, which is stable, and treat ticker, isin and the other identifiers as attributes of a row. Map them to your security master as of each decision date, not as of today, so a ticker that changed hands never pulls another company's history into the test.
Some datasets describe a state that moves: Technology Stack, with last_seen_at and missing_since, the listed securities table with status, and the latest observation in Website Profile. A load of such a dataset shows its state on the day you took it, so load those on a schedule and keep every load. Where a dated change dataset exists, such as Technology Changes for technologies, use it to fill the days between loads.
Querying as of a date
Keep a table of decisions, one row per company and date you will trade or score, and put the knowledge rule in the join condition, where it cannot be forgotten. The query below sums new and closed postings from Hiring Activity over the 28 closed days that were knowable, with one day of lag, before each decision. It leaves reconstructed rows out.
select
d.isin,
d.decision_date,
sum(h.new_postings) as new_28d,
sum(h.closed_postings) as closed_28d,
count(distinct h.day) as days_found
from decisions d
join company_hiring_daily h
on h.isin = d.isin
and h.day > d.decision_date - 30 -- 28 days, ending two days before the decision
and h.day <= d.decision_date - 2 -- closed, plus one day of lag
and not h.reconstructed
group by d.isin, d.decision_date;For a dataset that describes a state, take the last full load at or before the decision. This assumes you load the whole table each time.
select domain, technology, first_seen_at, last_seen_at, missing_since
from company_technologies_loads
where loaded_at = (
select max(loaded_at)
from company_technologies_loads
where loaded_at <= timestamp '2026-10-01 00:00:00'
);Report what the rules removed: rows left out as reconstructed, decisions with fewer than 28 days found, and the share of decisions with no data at all. A test that reports them states its coverage exactly.
Then run the test at several lags, for example one, two and five days, and plot the result against the lag. A signal whose value disappears after a day or two is a speed advantage you may not have. A signal that decays slowly is one you can still trade with a realistic delay.
Three checks before you trust a series
- Pull a closed period twice. Fetch the same closed week a week apart and compare the rows. Daily and weekly rows are never changed, so the two pulls match row for row. Datasets that describe a state differ by design.
- Count reconstructed rows. Tabulate the flag by dataset and by week. A test that excludes those rows should not lose a material share of its observations, and if it does, say so beside the result.
- Measure the posting lag. For each job board, compare
posted_atwithfirst_seen_at. A board that dates postings well before they were first seen will look early if you use the wrong column.
A list built for each date
Build the list for a date from what existed on it. The listed securities table keeps a row for every listing. A listing that has left the market is marked status delisted and stays in the table, so an old holding still finds its company. Keep the delisting date in your security master; the guide to mapping alternative data to a security master shows the join.
Coverage needs the same care. The company index grows, and a company enters a series on the day of its baseline observation. Take its first row in the dataset you are testing as its entry date, and admit it to a signal only when the signal's look-back window lies wholly after that date. The weekly market series apply the same rule through same-store cohorts.
Reading versions
Write-once guards against revision, and a version number guards against redefinition. A posting keeps the label version it was given until it closes, so a series that spans a version change mixes two definitions. Every jobs-v1 key keeps its meaning in jobs-v2, so restrict a series to those keys or to one version. The methodology documents each version and its change notice, and the data dictionary lists the columns that carry the label version.
Frequently asked questions
What is point-in-time data?
Point-in-time data records each value as it was known on its date and does not change it afterwards. A back-test on it can use only information that existed on the day of each decision. Every observation is dated and written once, so the history it shows is the history an investor could have seen.
How do I avoid look-ahead bias with alternative data?
Date every row by when it could be known, not by when the event happened, and add a lag you can defend. Read only rows whose knowledge date precedes the decision, keep your own load times, leave out periods written late and use a vendor that writes each period once and never revises it. Then check that the result survives a longer lag.
What does the reconstructed flag mean on daily and weekly rows?
It marks a daily or weekly row whose period was written more than seven days after it closed. The row is still written once and never changed, but it was not available on the normal schedule. A strict back-test leaves such rows out, or treats them as known only from the day you first received them.
Why is the first observation of a website or job board not counted as a change?
The first observation of a website or job board shows everything on it, and none of it appeared that day. Fokals stores that observation as a baseline, writes no change events for it and leaves first-observed postings out of the count of new postings, so a company entering the data does not look like a burst of activity.
How should I store vendor data so I can query it as of a date?
Append every load untouched with your own load time, key each dataset on its natural key, and derive a knowledge date for each from its date column plus a lag. Read research data only through a view that filters on the knowledge date. Keep every load of any dataset that describes a state that changes.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.