Library
Glossary
The terms of company data, defined: firmographics, technographics, intent, identifiers, licensing and delivery.
A
- Account scoringAccount scoring ranks companies by fit and timing so that sales and marketing know where to start. How a score is built and tested, and which inputs Fokals supplies.
- Account-based marketing (ABM)ABM starts with a list of target companies and builds campaigns around each. How a programme runs, and the company data that builds the list and times each account.
- Alternative dataJob postings, web traffic, card transactions and other data from outside financial statements and prices. How investors use it and the five checks to run on any source.
- Applicant tracking system (ATS)An applicant tracking system is where employers publish jobs and manage candidates. This entry covers what it means for hiring data, how it shows in hiring data and a common trap.
B
C
- Central Index Key (CIK)The number the US SEC gives each company, fund or person that files on EDGAR. What it identifies, how to format it and how to use it to join filings to a company.
- Company identifierA company identifier is a code that stays with a company when its name changes. This entry sets out what each common identifier identifies and how to choose a key.
- Corporate hierarchyA corporate hierarchy is the ownership structure of a company group. This entry covers parents, subsidiaries and brands, how to roll data up, and how Fokals links brands to listed parents.
- Cursor paginationCursor pagination marks where a page ended so the next call can resume from there. This entry compares it with offset paging and shows how to use it safely.
D
- Data coverageData coverage is how many of the companies you track a dataset contains and how fully each record is filled in. How to measure it on your own list, and how Fokals states coverage dataset by dataset.
- Data dictionaryA data dictionary says what every table and column of a dataset means. This entry lists what a useful one states and how to test it against a sample.
- Data enrichmentData enrichment adds fields from an external source to records you already hold. This entry shows the steps, a table of fields with refresh rates and common mistakes.
- Data freshnessData freshness is the age of a record's observation when you use it. How it differs from refresh cadence, how to measure it in your own warehouse, and the cadence of each Fokals dataset.
- Data lakehouseA data lakehouse puts a table layer over open files in object storage. This entry explains the parts, how vendor files are landed, and why time travel is not point-in-time data.
- Data licenceThe contract that gives a buyer the right to use a dataset, and sets its limits. Internal use, embedding and redistribution, the terms to settle and how terms travel.
- Data manifestA data manifest is the file that describes one delivery of data. This entry shows what it holds, how to check a load against it and what Fokals states in its own.
- Data marketplaceA data marketplace is a venue where providers list datasets and buyers acquire them. This entry separates its two kinds and lists what to check beyond the listing.
- Data provenanceData provenance records where data came from, when it was observed and what was done to it. What a record should hold, and how source, observation time and label version appear in Fokals tables.
- Data sharingData sharing gives another organisation read access to tables you hold without sending a copy. This entry explains how it works and what it means for licence terms.
- Data warehouseA data warehouse stores integrated, historical data modelled for analysis. This entry explains grain and keys and shows how vendor files fit a warehouse model.
- Domain matchingDomain matching compares the website domain two records carry. This entry shows how to normalise domains, where the method fails and how to build a domain-to-company map.
- Due diligence questionnaire (DDQ)The questionnaire a buyer sends a data vendor to check sourcing, rights, personal data and MNPI. What it asks, which standard forms exist and how to answer with documents.
E
F
- FIGIThe open 12-character code that names a financial instrument at three levels: one listing, a country composite and a global share class. How it differs from an ISIN.
- Firmographic dataThe attributes that describe a company: industry, size, location, ownership and growth stage. How teams use them to segment a market, and what to check in a source.
- First-party dataFirst-party data is what an organisation collects directly from its own customers and operations. Two senses of the term, and why the second matters in due diligence on a data vendor.
- Form 8-KThe current report a US company files within four business days of a significant event. The numbered items, a table of the common ones and how to read them in a dataset.
- Form DThe notice a company or fund files within 15 days of the first sale in an offering exempt from registration. What the form shows, how it is used and where it stops.
H
I
- Ideal customer profile (ICP)An ICP describes the companies that gain most from a product and return most to the seller. How to build one from won accounts, test it against losses and keep it honest.
- Incremental syncIncremental sync fetches only the records added since the last run. This entry shows the bookmark, the load and the failure modes, with an example on daily company data.
- Installed baseAn installed base is the set of customers using a product today. How to estimate it from outside with technology detection, follow its movement week by week and read the count honestly.
- Intent dataSignals that a company is researching or about to buy, usually scored by topic. The kinds of intent data, how to read a score and how an evidence-based score is built.
- Intent surgeAn intent surge is a sharp rise in a company's interest in a topic above its own normal level. How the flag works, a surge and a near miss worked through, and how to read the evidence.
- ISINThe 12-character international code for a share or a bond: a country prefix, a national number and a check digit. What it identifies, how to validate it and where it stops.
J
- Job functionA job function is the kind of work a role does, whatever its title. This entry shows how it is labelled, queried and read, and where function labels are uncertain.
- JSON LinesJSON Lines holds one JSON record per line, so large exports can be streamed and appended to. This entry gives the rules, a worked example and the traps.
L
- Label versionA label version names the frozen rules and model behind a dataset's labels. Why it matters for comparing periods and for products built on labels, and the versions used in Fokals data.
- Lead scoringLead scoring ranks people in your funnel by fit and behaviour. How points work, how company-level data can improve the fit side, and how to keep a score from going stale.
- LEIThe 20-character code for a legal entity such as a company or fund. What a record holds, how the check digits work and how an LEI maps to the securities it issues.
- Look-ahead biasA back-test that uses information from after the date it simulates looks better than any real decision could have been. How it enters a test and how to find it.
M
- Market identifier code (MIC)The four-character ISO 10383 code that names an exchange or trading venue. Operating and segment MICs, why a ticker needs one and the join mistakes to avoid.
- Material non-public information (MNPI)Information that a reasonable investor would find important and that has not been released. The two tests, the questions a fund asks of a data source and common mistakes.
O
P
R
- Rate limitA rate limit caps the requests a client may send in a period. This entry shows how to plan a backfill against minute and daily limits and how to act after a refusal.
- Redistribution licenceA redistribution licence lets a buyer pass licensed data on to its own customers. This entry sets it beside internal use and embedding and lists the terms to settle first.
- robots.txtrobots.txt is the file at the root of a website that tells crawlers which paths to avoid. How its rules are matched, what happens when it is missing, and why it is not security.
S
- Same-store cohortCompare only the companies present for the whole window, so wider coverage does not look like market growth. The method, a worked example and where it is used.
- Security masterThe reference table that gives every security one identity, so prices, holdings and alternative data join without guesswork. What it holds and how to map a dataset to it.
- Seniority levelSeniority level places a job on a ladder from intern to executive. How it is read from a posting, where titles mislead, and how it appears in Fokals hiring data.
- Survivorship biasStudying only the companies that lasted makes past results look better than they were. Where the error appears, how to test for it and how to keep delisted names.
T
- Tag managerA tag manager lets a site owner change scripts from one interface. This entry explains how containers work and why reading them gives a complete view of a site's technologies.
- Technographic dataThe analytics, advertising, commerce and CRM tools a company runs, found by matching signatures in public code. How detection works and how to read it.
- Technology churnTechnology churn is the share of companies that stop using a technology in a period. How to calculate it, read removals from outside without false alarms, and spot a switch.
- Third-party dataThird-party data is information collected by someone else and supplied to you under licence. What to ask a supplier before you rely on it, and how company-level data differs from audience data.
- Total addressable market (TAM)TAM is the yearly revenue available if every company that could buy did buy. How to build it from a count of companies, and how company data supplies the count.
- Traffic rankA traffic rank orders websites by how much traffic they receive. How rank buckets work, what a rank can and cannot say about a company, and the monthly traffic tier in Fokals data.
- Trigger eventA trigger event is a dated change at a company that gives a reason to reach out now. How triggers are paired with plays, checked for date and duplicates, and read from company announcements.