Glossary

Entity resolution

Entity resolution decides which records describe the same company and links them under one key. This entry covers the order of matching, how to measure it and the rules Fokals publishes.

Updated 5 October 20262 min read

Entity resolution decides which records, in one dataset or across several, describe the same real-world entity, such as a company, and links them under one key. It uses identifiers and normalised attributes such as names and domains, and it is judged by how many links are wrong and how many are missed.

How matching proceeds

A sound process works from the strongest evidence to the weakest. First normalise: fold case, strip punctuation and legal suffixes, and standardise country names. Then match on shared company identifiers, which are exact and cheap. Next match on a normalised website domain, as described under domain matching. Last, compare names together with country, and send uncertain pairs for review instead of accepting them. Links are then grouped into entities, and each link keeps the rule that made it.

An illustration with Acme Robotics: the names Acme Robotics Inc. and ACME ROBOTICS, INC both normalise to acme robotics. A different company with the same name in another country must not join them, and the country is what separates the two.

Measuring it

Two figures judge a match. Precision is the share of links that are correct, and recall is the share of true matches that were found. The match rate, the share of records that found a partner, measures neither: a rule that links everything has a perfect match rate and poor precision. Measure both on a sample checked by hand.

Suppose a rule makes 100 links in a sample and a hand check finds 96 of them correct: precision is 96 percent. If the sample holds 120 true matches, recall is 96 out of 120, or 80 percent. In finance a wrong link can cost more than a missing one, because it assigns one company's hiring or filings to another company's ticker, so rules tend to be conservative.

In Fokals data

The methodology states the rules:

  • Every listed security is tied to its company by identifier first (ISIN, LEI), then by ticker on the same market, then by exact normalised name in the same country. Each listing records how it was matched, so a link can be audited.
  • Funding matching is exact. A private capital raise in Company Funding is tied to a company when the issuer's normalised legal name matches one company in the index, which keeps the links precise.
  • An ISIN or LEI is stored only when its check digit validates.

The data dictionary describes the fields, and the use case on mapping alternative data to a security master puts the join to work.

Frequently asked questions

What is the difference between entity resolution and record linkage?

In practice the terms overlap. Record linkage usually means matching records across two datasets, as when a customer file is joined to a register. Entity resolution is the broader process: it also covers finding duplicates within one dataset and grouping all the records of one entity under a single key.

How do I match company names across datasets?

Do not match on names alone. Match on a shared identifier first, then on a normalised domain, and use name with country only for what is left. Normalise case, punctuation and legal suffixes such as Inc or GmbH before comparing, and send close but uncertain pairs for review instead of accepting them.

How accurate is entity resolution?

It depends on the keys available and on how cautious the rules are. Judge it with a sample checked by hand: precision is the share of links that are right, and recall is the share of true matches found. A rule that matches every record has a perfect match rate and poor precision, so ask for both figures and not the match rate alone.