Domain matching joins records that describe the same company by comparing the website domain each carries, after reducing every domain to one standard form. It is fast and usually precise, and it fails where a company runs several domains or a domain is shared.
Normalising a domain
Both sides of the join must be reduced in the same way. The steps are:
- Lower-case the value and trim it.
- For a web address, drop the scheme, port, path and query. For an email address, keep the part after the
@. - Remove a leading
www.. - Reduce the host to its registrable domain with the Public Suffix List, so that
shop.acme-robotics.examplebecomesacme-robotics.exampleandacme.co.ukis not cut toco.uk. - Convert internationalised names to one form, such as punycode.
Illustrative inputs for Acme Robotics:
| The record holds | Normalised |
|---|---|
https://www.Acme-Robotics.example/about | acme-robotics.example |
sales@acme-robotics.example | acme-robotics.example |
shop.acme-robotics.example | acme-robotics.example |
Where it fails
- A company with several domains, such as a brand site, a country site or an acquired company's domain, matches only the one you hold.
- A redirected or retired domain matches nothing, or the wrong company.
- A shared domain, such as a page on a marketplace or a free email service, ties unrelated records together.
- A domain can change hands, so an old record may point to a new owner.
- A brand's domain matches the brand and not its parent, and the corporate hierarchy decides what to do next.
In Fokals data
Technology Stack and Web Traffic each carry the website domain of the company, and every dataset carries the Fokals company ID. A domain-to-company map is therefore one query, after which every other dataset joins on the company ID:
select a.account_id, m.company_id
from crm_accounts a
join (select distinct domain, company_id from company_technologies) m
on m.domain = a.normalised_domain;Here the CRM accounts table is your own. Each company in the index carries the domains of its own websites, so the map is built on the company's published domain and never on a guess. The data dictionary lists the fields, and the marketing stack dataset holds Technology Stack and its companion datasets.
Related terms
- Entity resolution: the wider process of deciding which records are one company.
- Company identifier: the codes that match more exactly than a domain can.
- Data enrichment: the step that follows a successful match.
Frequently asked questions
How do I match companies by website domain?
Normalise the domain on both sides: lower-case it, strip the scheme, path and leading www, and reduce it to its registrable domain. Join on that value, then review the cases where one domain matches several records. Match the leftovers by identifier, or by name with country, and keep a note of how each match was made.
Why does domain matching miss some companies?
Usually because a company uses more than one domain, the record holds an old or redirected domain, or the domain is shared, as with a page on a marketplace. A company that has no public website needs another key, such as an identifier. Check a sample of non-matches by hand to see which cause dominates.
Should I match on domain or on company name?
Domain first. A normalised domain is far more specific than a name, which repeats across companies and varies in spelling and legal suffix. Use names only for records with no domain, and require a second field, such as country, before accepting a match. Keep how each match was made, so that you can review the weaker ones.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.