Choosing a vendor of company data comes down to asking every candidate the same questions and declining answers you cannot check. This post gives you ten, in the order a review usually runs: what the vendor holds, whether you can use it, and what the contract allows. Each comes with the evidence to ask for and a test to run on sample data. We then answer all ten ourselves, with the public document behind each answer.
Before you ask
Three preparations make the answers comparable between vendors.
- Write down the companies you track: the ones the data must cover, as a list of domains or security identifiers.
- Write down the decision the data feeds and how often it is made. Freshness and the dating of the record are judged against that.
- Ask each vendor for sample data on the companies you track, not on a list of its choosing, together with its data dictionary.
Then grade every answer by the evidence behind it: an assurance given in a call, a published document, or a result you reproduced on the sample. Count only the last two.
What the vendor holds
1. Coverage: how many of the companies I track do you hold, field by field? A total count says how large the vendor's index is. Data coverage is the share of your list it matches and, among the matches, the share with each field you need. Ask for both by country and by company size. Test: match your list to the sample and compute the two rates yourself.
2. Sourcing: where does each record come from, and how was it gathered? Ask for a written sourcing statement. It should say what kinds of source the data comes from, whether anything sits behind a login or a paywall, how the vendor's collection identifies itself to the sites it visits, and whether personal data is involved. Test: trace twenty sampled records back to the original page or filing.
3. Freshness: how old is each observation when it reaches me? A file delivered daily can carry records observed a month ago. Ask for the refresh cadence of each dataset, and for the observation time on every record, which is what data freshness measures. Test: compute the age of each sampled observation and read the oldest tenth, not the median.
4. Observation: was each value observed on its date, or rebuilt afterwards? A record observed on its date shows what could be known that day. A record assembled later shows what is known now, and a test run on it acts on information nobody had at the time. Ask which records carry an observation time, and how the first observation of a company is kept apart from a change. Test: take twenty dated changes from the sample and confirm that each carries the time it was observed as well as the date of the event.
Whether you can use it
5. Identifiers: what do I join on? Ask for a stable company key, the website domain and, for listed companies, the security identifiers your systems use. Ask how a brand or subsidiary is tied to its parent. Our note on the identifiers that make company data joinable sets out the levels. Test: join the sample to your CRM or security master and count the matches by key.
6. Point-in-time: is a closed period ever rewritten? If rows change after they are first published, a back-test runs on values nobody had on the day. Ask whether closed periods are revised and how a period written late is marked. A vendor that answers never should be able to show you how to verify it. Test: pull the same past period twice, some weeks apart, and compare. The entry on point-in-time data explains the design.
7. Labels: which values were observed and which were inferred? An industry, a job function or an intent score is the output of a model. Ask which fields are model output, which label version produced each row, how confident the model must be before a label is written, and how accuracy is measured. Test: grade fifty sampled labels by hand against the evidence the model read.
8. Delivery: how does the data reach my systems? Ask which routes and formats are offered, how an incremental load finds where it stopped and what the rate limits are. Test: load the sample into your warehouse and run the incremental step twice. The second run should add nothing.
What the contract allows
9. Licence: what may I do with the data? Internal use, embedding in a product and redistribution are three different rights. Ask which one the quoted licence grants, what it says about derived data such as scores and models, and what happens to the data you hold when the term ends. The three are set side by side in what a data licence covers.
10. Change management: what happens when a definition changes? Ask how a change to a label, a score or a schema is released: under a new version name or in place, with how much notice, and whether old and new are delivered side by side. A change made in place alters every model built on the old meaning, and nothing in the data shows it.
How we answer the ten
Here are the same questions, put to Fokals. The documents behind the answers are public: the methodology, the data dictionary, the sourcing statement and the API reference.
| Question | Our answer |
|---|---|
| 1. Coverage | Listed companies worldwide, with equity listings in 79 countries, the brands they own and verified private companies, on one company index. Sample data for the companies you track is sent on request, with the data dictionary and the methodology. |
| 2. Sourcing | First-party company sources and public records, processed in-house: what companies publish on their own websites and careers pages, what they announce and what they file. Company-level data throughout. The sourcing statement is public. |
| 3. Freshness | Hiring daily; marketing stack daily to weekly; intent scores weekly and signals daily; announcements and scale daily to every three days; Market Series weekly, as of each Sunday. Every record carries the time it was observed. |
| 4. Observation | Observed, and dated as observed. Every record carries its observation time. A first observation sets a baseline and is never counted as a change. |
| 5. Identifiers | One stable company ID. For a listed company: ticker, MIC, ISIN, LEI and share-class FIGI, and the SEC CIK where it has one. Brands and subsidiaries carry their listed parent's identifiers. |
| 6. Point-in-time | Each daily and weekly dataset is written once, after its period closes, and is never revised. A period written late carries a flag. The record is point-in-time by construction, and section 10 of the methodology sets out the time conventions. |
| 7. Labels | Model output is produced under named, frozen versions, and the version is carried on the rows. A label is written only when the model clears a fixed probability floor, and accuracy is measured by hand for each version. |
| 8. Delivery | A REST API of 25 endpoints, with one bearer key per client, rate limits per key and cursor pagination, and bulk exports of any dataset for any period as JSON, JSON Lines or CSV, with a manifest. The feeds of signals, website changes and postings run oldest first from a time you set, so the last cursor is the bookmark for incremental sync. |
| 9. Licence | A written agreement for one or more of three uses: internal use, embedding in a product, redistribution. |
| 10. Change management | Labels and scores are frozen under named versions, and a breaking change ships as a new version with at least 90 days' notice. Additions to the technology catalogue are recorded in the changelog of the methodology. |
Every one of these answers can be tested. Sample data for the companies you track comes with the data dictionary and the methodology, so the ten tests above run on your list of companies and against the documents that define each field.
Three of the answers work together for research use. The observation time on every record, the baseline rule and the write-once rule are what make the record point-in-time, and why we write each table once sets out how to check each of them on two loads of the same week.
What the ten questions do not settle
The questions test whether a dataset is what its vendor says it is. Whether it helps your decision is a separate test, and it takes your own outcomes: set the sample beside the thing you want to predict or explain, and see whether the data moves first.
One question also comes before all ten: whether the vendor holds the kind of data your decision needs. Fokals is company-level data throughout: firmographic, technographic, hiring and intent data on public and private companies, joined on one company ID. The datasets page lists what each product delivers.
Frequently asked questions
What questions should I ask a data vendor before buying?
Ask ten: how many of the companies you track the vendor covers, where each record comes from, how old each observation is, whether each value was observed on its date or rebuilt later, which identifiers you can join on, whether closed periods are ever rewritten, which values are model output, how the data is delivered, what the licence allows, and how a change of definition is released. For each, ask for a document and run a test on sample data.
What are the warning signs when evaluating a data vendor?
Five recur: a total record count offered in place of coverage on your own list; records with no observation time, so that freshness cannot be measured; past periods with no account of whether they were observed at the time or rebuilt afterwards; labels or scores with no version name, so that a change cannot be seen; and definitions changed in place, with no notice. Each one leaves you unable to check a claim the vendor has made.
What documents should a company data vendor provide?
A vendor should provide four documents: a sourcing statement that says where the data comes from and how it is gathered, a methodology that says how it is labelled and scored, a data dictionary that defines every dataset and field, and a reference for the delivery method. Fokals publishes all four, as the sourcing statement, the methodology, the data dictionary and the API reference. Licence terms are in the written agreement.
How do I compare two data vendors fairly?
Give both the same list of companies, the same questions and the same tests, and grade the evidence, not the confidence of the answer. Compare coverage on your list, never on total record counts, because a larger index can still hold fewer of the companies you track. Compare freshness by the age of observations, and the record by what was observed at the time, not by what was rebuilt later.
What makes company data safe to use in a back-test?
Three properties, each of which you can test on a sample. Every record carries the time it was observed, so a query as of a past day returns what was known that day. A first observation is a baseline and is never counted as a change. And closed periods are written once and never revised. Fokals datasets are built on all three, which makes the record point-in-time by construction.