Look-ahead bias is an error in a back-test or model that uses information not yet available on the date being simulated. The result looks better than any real decision could have achieved, because the test, in effect, knew the future.
How it gets in
- Restated values. A figure is corrected later and the corrected value is used for the original date.
- Publication lag. A number is dated by the period it describes, not the day it was published.
- Backfilled history. A vendor adds years of history that never existed as data on those days.
- A hindsight list. The test starts from today's index members, which is also survivorship bias.
- Later labels. A label or model built with data from after the date is applied to that date.
- Full-sample statistics. A mean or a threshold is computed over the whole period, then used inside it.
How to find it
For each field, ask when its value first existed and could be acted on. Join on that moment, and add the delay you would have faced in practice. Compare two downloads of the same past period: any value that changed shows where the history moves. Then run the rule on a later period it has never seen. If the result falls away, an input was looking ahead. The use case for back-tests shows how to store and query data as of a date.
In Fokals data
The Job Postings dataset gives every posting its first-seen date, the time it was first observed, beside the date the employer published it. Use the first-seen date as the date a posting was known. The published date can be earlier, and a posting already open when first observed is flagged, because its first-seen date then falls after it opened.
Every observation is dated. Daily and weekly tables are written once, after the period closes, and never revised. A flag marks any period written more than seven days after it closed, so a test can leave it out. Labels are produced under frozen label versions, and a test is split where a version changes. The data dictionary defines each field.
Related terms
- Point-in-time data: data stored as it stood on each date.
- Survivorship bias: testing only on the companies that lasted.
- Label version: the frozen name under which a label is produced.
- Alternative data: the family of sources most often tested on history.
Frequently asked questions
What is an example of look-ahead bias?
A back-test buys each stock on the last day of a quarter using that quarter's reported earnings. The earnings were published weeks later, so no investor could have traded on them that day. The strategy appears to predict returns because it already holds the answer. Dating each figure by its publication day, plus a realistic trading delay, removes the error.
How do you avoid look-ahead bias in a back-test?
Use point-in-time data, stored with the date each value became known. Join on that date, add the delay you would face in practice, rebuild the list as it stood on each date, and fit any model only on data from before the date it is applied to. Then confirm the result on a later period the rule has never seen.
Is look-ahead bias the same as survivorship bias?
No, though they often occur together. Look-ahead bias uses information from after the test date. Survivorship bias keeps only the entities that still exist at the end, so the failures are missing. A test on today's index members has both: it knows who survived, which nobody knew then, and it omits those that did not.