Glossary

Data freshness

Data freshness is the age of a record's observation when you use it. How it differs from refresh cadence, how to measure it in your own warehouse, and the cadence of each Fokals dataset.

Updated 5 October 20262 min read

Data freshness is how recent a dataset's information is at the moment you use it: the gap between when an observation was made and now. It differs from how often a dataset is refreshed, because a feed that runs daily can still carry records last observed days or weeks earlier.

Four clocks

Each record passes four moments, and a delay at any of them looks the same to the person reading it. The event happens in the world. The vendor observes it. The vendor publishes it in a dataset. You load it. Freshness is the age of the observation when you use it; latency is the whole delay from event to load. Measure the clocks separately and you can tell a slow source from a slow pipeline.

Measuring it

Measure the age of each record's observation, not of the file that carries it, and look at the distribution: a median of one day can hide a tail of a month. Set the threshold by the decision the data feeds, since a weekly score is not late after three days. In a table of website observations, the age of each company's latest reading is:

select
  company_id,
  now() - max(last_read_at) as age
from company_site_facts
group by company_id
order by age desc;

Freshness in Fokals data

Cadence differs by dataset. Technology Stack and Website Profile refresh daily to weekly, Job Postings and Hiring Activity daily, Company Signals daily and Intent Scores weekly, Company News daily to every three days, and Web Traffic monthly. Each record carries the time it was observed, so the age of every value can be measured.

Daily and weekly rows are written once, after the period closes, and are never revised, which keeps them point-in-time data. The data is refreshed daily and delivered by REST API and bulk files, so a copy you hold is as fresh as your last pull. The methodology gives the cadence rules, and delivery describes how the data reaches you.

Freshness is read beside data coverage and data provenance. Tables written once after a period closes are point-in-time data, and a daily pull is an incremental sync.

Frequently asked questions

What is the difference between data freshness and data latency?

Freshness is the age of the data at the moment you use it, measured from when the observation was made. Latency is the delay between an event in the world and the moment the data about it becomes available to you. They are related: latency sets the best freshness a pipeline can reach, and the time since the last refresh adds to it. Measure both, because they point to different fixes.

How fresh does company data need to be?

It depends on the decision. Prioritising accounts for a weekly review needs weekly data. Reacting to a hiring surge or a leadership change calls for daily. A quarterly market model can use monthly. Match the cadence to how quickly the decision loses value.

How do you find the age of a record?

Read the observation time on the record and subtract it from now, as in the query above, then look at the distribution across the companies you track. Every Fokals record carries the time it was observed, so the age is measurable for every value and not only for the file that carries it. Compare the result with the threshold the decision needs.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.