We write each daily and weekly dataset once, after the period closes, and it is never revised. This post explains why, and sets out the four rules that go with that one: every observation carries its time, a first observation is a baseline and never a change, a late period is flagged, and labels are frozen under named versions. If you back-test, train models or build a product on dated company signals, it tells you what these rules guarantee and how to check them on the data yourself.
What a rewrite does to a table
A dataset can be kept in two ways. It can show the best present knowledge of the past, corrected whenever something new is learned, or it can show what was known on each date and leave it there. The first suits a chart. The second is point-in-time data, and it is what a back-test needs, because a rule tested on corrected history acts on values nobody had on the day. That error is look-ahead bias.
Corrections are not the only way a history misleads. It also misleads when an event is dated by when it happened and not by when it could be seen, when a newly covered company arrives with everything on its site counted as new, when a period is filled in weeks later without a mark, and when a label keeps its name while its definition changes. Each of the five rules answers one of these.
The five rules
| Rule | What it means | Where you see it |
|---|---|---|
| Written once | A daily or weekly row is written after its period closes and never rewritten | The day in Hiring Activity, the week in Sales Team Metrics and Intent Scores, the as-of date in Market Series |
| Observation times | Website changes, postings and signals carry the time they were observed | The observation time on changes and signals; the first-seen, last-seen and close dates on postings |
| Baselines | The first observation of a website or of a company's postings is a baseline and never a change | No change event for a first observation, and a baseline flag on postings |
| Late periods flagged | A period written after its normal schedule is marked | The reconstructed flag on the row |
| Frozen versions | Labels and scores come from named versions that do not change | The label version and the intent version on the row |
Written once, after the period closes
All times are UTC. Each closed UTC day gets its rows in Hiring Activity once. Each closed week, Monday to Sunday, gets its rows in Sales Team Metrics and Intent Scores once. Market Series is written once for each as-of date, a Sunday, so whenever you read a row it holds what was known on its Sunday.
Every observation carries its time
Not everything is a period. A change on a website, a posting and a signal are events, and each carries the time it was observed: the observation time in Technology Changes and Company Signals, and the first-seen, last-seen and close dates in Job Postings.
An observation time is not always the date of the event, and the datasets keep the two apart. In Job Postings the posting date is the date the employer gives, and the first-seen date is when the posting was first observed. In Company News the date is the day the company published the announcement, or the time it was first observed where the announcement itself carries no date. Choose the date that says when you could have known.
A first observation is a baseline
A first observation records what is already in place: the stack a website runs, the roles a company has open. It sets a baseline and writes no change, so every row of Technology Changes is a change observed between two dated observations of the same website. Job Postings works the same way: a posting already published when a company's postings are first observed carries a baseline flag and is not counted among the new postings of Hiring Activity. New listings follow the same rule, and the first observation of a market sets its baseline.
With this rule a company added to the index brings its present state and no false burst of activity. The weekly Market Series go a step further for the same reason and take every rate over a same-store cohort.
A late period is flagged
A period written after its normal schedule carries the reconstructed flag. The row is written once like any other, and the flag tells a strict test that it arrived late, so the test can leave it out.
Labels and scores are frozen under a name
A label made by a model changes when the model, the question or the threshold changes. So each set of labels has a version name, and the labels, the questions, the model and the probability floors behind that name are fixed. Any change comes out as a new label version.
The same holds for the intent score: a change to topics, weights or the scoring formula is a new intent version. A breaking change is released with at least 90 days' notice. When the definition of a Market Series metric changes, the metric gets a new identifier and the rows already written stay as they are.
The version travels with the row. A posting keeps the version it was labelled under, so the label version in Job Postings always says which definition produced a value.
What the rules guarantee
Taken together the five rules give a reader four things to rely on.
- A past day reads as it was written. A query as of any past day returns what was known that day, with nothing learned later mixed in.
- Growth in coverage stays out of the signal. A company added to the index sets baselines and adds no changes on arrival, so a rise in a count is a rise in activity.
- A late period is marked as late. It is written with its flag and is never passed off as on time.
- A better definition leaves the record intact. It starts a new version or a new metric identifier and leaves the rows already written as they are, so a model trained on the old definition keeps its inputs.
Section 10 of the methodology states the time conventions in full.
How to check it yourself
Keep two loads of the same closed week, taken some weeks apart, and compare the rows that appear in both. Below, two loads of Intent Scores are stored as a first and a second table of your own, and the query should return nothing.
select a.company_id, a.topic,
a.score as first_load, b.score as second_load
from intent_load_1 a
join intent_load_2 b
on b.company_id = a.company_id
and b.topic = a.topic
and b.week_start = a.week_start
where a.week_start = date '2026-09-21'
and (a.score <> b.score or a.surge <> b.surge);The same comparison works on any daily or weekly dataset: join two loads on the company ID and the period, and compare each measure. An empty result is the write-once rule, verified on your own copies.
Then count the late periods, so you know how much a strict test would leave out.
select week_start,
count(*) as all_rows,
sum(case when reconstructed then 1 else 0 end) as late_rows
from company_intent_weekly
group by week_start
order by week_start;State datasets and the dated record beside them
The write-once rule applies to rows that describe a closed period. Some datasets describe a current state, and a state moves with each observation.
- In Technology Stack, the last-seen date of a technology advances with each observation.
- Website Profile holds the latest observation of each website.
- A record of Job Postings gains its close date when the posting closes.
- A listing that leaves its market is marked as delisted and kept, so joins to earlier periods still work.
- Website labels are produced again when a website changes or when the label version changes.
For each of these the dated record sits beside the state: Technology Changes for a website, Hiring Activity for a company's postings. If you need the state as it stood on a past date, load the state datasets on a schedule and keep each load. The guide to building a point-in-time dataset for back-tests shows how to store and query both, and the data dictionary defines every field named here.
Frequently asked questions
What is a write-once table?
A write-once table is one whose rows are written a single time and never updated afterwards. For period data, each day or week gets its rows after the period closes, and those rows stay as they were first written. A reader who loads the table a year later sees the values that were recorded at the time, which is what a back-test or an audit needs.
Why do data vendors revise historical data?
Usually to improve it: an error is corrected, a source arrives late, coverage grows and earlier periods are filled in, or a definition changes and history is recomputed to match. Each revision makes the present picture of the past better and the record of what was known at the time worse. Ask a vendor which of these it does, and whether revised rows are marked.
Can point-in-time data be corrected?
Not in place, or it stops being point-in-time. A correction has to arrive as something new and marked: a later row, a flag or a new version. Fokals writes daily and weekly rows once and leaves them as written, flags a period written after its normal schedule as reconstructed, and releases a changed definition under a new version or a new metric identifier.
Is write-once data the same as point-in-time data?
Write-once is one part of it. It guarantees that a period's rows do not change after they are written. Point-in-time use also needs a date for when each value could be known, a list of companies as it stood on each date, and labels whose definitions are fixed. A write-once table with undated events or shifting labels can still mislead a back-test.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.