Use case

Training and evaluating models on dated company signals

A model on company signals is only as honest as its feature table. This guide builds one as of each date, splits by time and names the properties of the record that keep it safe.

Updated 5 October 20267 min read

A model trained on company signals fails quietly when a feature held something that was not known on the date of its row. That is look-ahead bias, and this guide shows how to avoid it with Fokals datasets: how to build a feature table as of each date, how to split by time, how to treat frozen label versions, and which properties of the record keep a training set safe, namely baseline observations, write-once rows, versioned labels and a same-store population. Every observation is dated, every daily and weekly dataset is written once and never revised, and the record is point-in-time by construction, so the same method applies to any window you choose.

A feature table as of a date

For each company and each as-of date t, compute the features only from rows that were known by t, and compute the target only from rows after t. The datasets allow this: the row of a closed day or week is written once and never revised, and each record says when it was observed. You can rebuild any as-of table from stored rows instead of keeping snapshots. The guide to building a point-in-time dataset covers storage, and this one covers the model.

Use a Monday as the as-of date. The rule for every dataset is then the same: a daily row is usable if its day is before t, a weekly row if its week_start is at least seven days before t, and an event if its observed_at is before t. A weekly row is written after its week closes, so also require that your own load time is before t. In the query below, the datasets are Hiring Activity (company_hiring_daily), Technology Changes (company_tech_events) and Intent Scores (company_intent_weekly).

-- covered_companies: companies in the files since before the lookback window began
select
  c.company_id,
  date '2026-10-05' as as_of,
  (select h.open_postings from company_hiring_daily h
    where h.company_id = c.company_id and h.day = date '2026-10-05' - 1) as open_postings,
  (select sum(h.new_postings) from company_hiring_daily h
    where h.company_id = c.company_id
      and h.day between date '2026-10-05' - 28 and date '2026-10-05' - 1) as new_28d,
  (select count(*) from company_tech_events e
    where e.company_id = c.company_id and e.category = 'technology'
      and e.change = 'added'
      and e.observed_at >= date '2026-10-05' - 90
      and e.observed_at <  date '2026-10-05') as tools_added_90d,
  coalesce((select max(w.score) from company_intent_weekly w
    where w.company_id = c.company_id
      and w.week_start = date '2026-10-05' - 7), 0) as top_intent_score
from covered_companies c;

Two details in the query carry weight. Intent Scores are written for the topics where a company shows evidence, so for a company in the files an absent weekly row means a score of zero for that topic and reads as 0. A company whose hiring has no postings has no hiring row, and that is a missing value to encode apart from a zero. Encode the two cases differently.

Properties of the record that keep a model honest

Five properties of the datasets matter to a training set, and each comes with a rule for using it.

Baseline observations. A website's or job board's first observation serves as a baseline and produces no change. For a tool present at the first observation, the first-seen date in Technology Stack (company_technologies) is the date of that observation. For a posting marked found_on_first_read in Job Postings (job_postings), the first-seen date marks when Fokals first saw it. A feature such as months since adoption mixes real dates with baseline dates, and the mix depends on when the company entered the index. Take adoption dates from Technology Changes, and leave found_on_first_read postings out of opening features.

A growing index. Companies enter the index every day, and a company covered for a week has few events because it has been observed for a week. Restrict each as-of table to companies covered since before the lookback window began, the rule behind a same-store cohort in Market Series. Use the earliest first-seen date among a company's rows in Technology Stack as a proxy for the day it entered, and treat it as an approximation.

Reconstructed rows. A row flagged reconstructed=true was written more than a week after its period closed, so it holds values that were not on the record on time. Keep these rows out of training and evaluation, or evaluate them as a separate fold and report both.

Event dates and knowledge dates. In Company Signals (company_signals), observed_at is a posting's own date where it has one, so a signal can carry a date earlier than the day it reached the data. In Company News (company_news), at is the publication date, or the date first seen when the source gives none. Where the question is what you could have known, add your own loaded_at to every row you store and filter on it as well.

Label versions. A dataset can hold more than one label version. In Job Postings, label_version names the version under which a posting was labelled, and earlier versions carry fewer keys than the current one. A feature such as tool_snowflake is empty on rows labelled under a version that did not ask for it, which reads as false and means not asked. Build features from the keys the versions share, which mean the same in both, or restrict to the current version, and record the version set in the training run.

Splitting by time

Split on as-of dates, never on rows. Train on early dates, validate on later ones, and leave a gap between them equal to the length of the target window, so that no training target reaches into the evaluation period. If the target is that a company adds a technology in the next 28 days, the target window of a row runs from its as-of date to 28 days after, and no feature may include an event after the as-of date.

An illustration: number the weekly as-of dates 1 to 10 and use a target window of four weeks. The target of the row at date 4 ends before date 8 begins. Train on dates 1 to 4, skip 5 to 7, and test on 8. Move the origin forward and repeat, and report the mean over folds with its spread.

Keep a final window untouched until the last evaluation. If the model must work for companies it has not seen, hold out a set of companies as well as a period. The same company appears in many as-of rows, and those rows are correlated.

Frozen versions and the model

Every label and score is made under a named label version that stays frozen once released, so the inputs of a model are defined by the versions beneath them: jobs-v2 for posting labels, sales-v1 for sales roles, news-v1, sec-items-v1 and roles-v1 for announcements, signal-v2 and intent-v2 for the weekly scores. Record the set in the model card and in the training run.

A breaking change arrives as a new version announced at least 90 days ahead. Use that notice as a retraining schedule. List the features that depend on the changed version, collect weeks under it, and compare performance on the same as-of dates under both versions when both are delivered. A model should receive the new version's values in production only after it has been tested on them. Each bulk export comes with a manifest that lists its period, label versions and licence. Keep the manifest of every export that fed a run, and you can rebuild its table later. The blog post on writing each table once explains the design behind that stability.

Evaluating with discipline

  • Baselines. Beat last week's value and the company's own trailing average before claiming anything, and compare with the base rate across the same-store population.
  • Metrics by week. Compute precision at k or lift for each as-of week and report the mean and the spread. A pooled figure hides a collapse in one week.
  • Drift. Plot each feature's distribution by as-of week. A shift that follows growth of the index and not behaviour points to a leak or a missing same-store filter.
  • Uncertainty. Resample by week and give an interval around each metric, so a reader sees how stable it is.
  • Scores as features. Fokals delivers evidence-backed intent scores: each score carries the dated signals behind it, so a model reviewer can open the source. An intent score summarises the signals of the trailing 90 days, which makes it a feature. Take the target from a different source or window, and never derive a target from signals that sit inside the same 90 days.

How to read the labels

A model assigns the labels, under a frozen version, so the same version gives the same label to the same posting at any later date, and an empty value reads as not known. Treat label quality as a property of the version and measure it on a sample you check by hand. Targets such as revenue, valuations or deal outcomes come from your own sources, joined to Fokals on ISIN or FIGI for listed companies and on the company ID for the rest. A population drawn from websites the datasets observe is the population your model will be asked about, so it is the right one to evaluate on.

How Fokals delivers it

The features above come from the hiring dataset, the intent dataset and the marketing stack datasets, delivered by REST API or as bulk exports. The methodology describes each label version and how each dataset is written.

Frequently asked questions

What is data leakage in a machine learning model?

Leakage is when a model's features contain information that would not have been available at the time of the prediction, so its test results overstate what it can do live. With dated company data the usual sources are baseline observations counted as events, rows written late, event dates used as knowledge dates, and target windows that overlap the features.

How do I split time-dependent company data for training and testing?

Split on as-of dates, not rows. Train on earlier dates, test on later ones, and leave a gap equal to the length of the target window so that no training target reaches into the test period. Repeat with the origin moved forward, and report the average and the spread across folds.

What makes company signals safe to train on?

Every observation carries the time it was observed, daily and weekly datasets are written once after the period closes and never revised, and a first observation sets a baseline and is never counted as a change. Labels and scores come under named, frozen versions. Together these make a feature table rebuildable as of any date, which is what a model needs.

Should I use intent scores as features or as labels?

As features. Each score carries the dated signals behind it, summarised over the trailing 90 days, so it is evidence a model can use as input. Derive the target from another source or window, because a target taken from the same signals would share evidence with the feature and flatter the test.

What happens to my model when a label version changes?

A breaking change comes as a new version announced at least 90 days ahead, so there is time to plan. Record the version set with every training run, test the model on the new version's values before production sees them, and retrain once enough weeks exist under the new version. Do not mix versions in one feature without checking that its meaning is unchanged.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.