All documentsMethod · Version 2.1

Methodology

Sources, collection, coverage, labelling and scoring for every dataset.

This document describes how the Fokals company-signals products are built: where the data comes from, how it is collected, how it is normalised and scored, and what its limits are. Field-level definitions are in DATA_DICTIONARY.md. How the data is sourced, and what is never collected, is in SOURCING.md.

1. Products

ProductWhat it answersMain filesRefresh
Marketing stackWhich advertising, analytics, commerce, marketing and sales technologies a company runs on its own website; what the site offers (markets, languages, currencies, apps, key pages, promotions); every change, dated.company_technologies, company_tech_events, company_site_factsDaily to weekly per company (section 4)
HiringWhat a company is hiring for: every posting on its job board with place, work mode, pay, requirements, tools named and labels; daily totals per company; the sales organisation: roles, segments, advertised base and on-target earnings, quota, ramp and lead mix, weekly, with pay benchmarks across companies.job_postings, company_hiring_daily, company_sales_weekly, sales_pay_benchmarksDaily
IntentWhat a company is likely to buy or build soon, by topic, with the evidence.company_intent_weekly, company_signals, company_fundingWeekly (signals daily)
Announcements and scaleWhat a company announces about itself (launches, partnerships, acquisitions, funding, leadership changes, layoffs, expansion, incidents, results), from its own newsroom and feeds and, for listed companies, its Form 8-K; its headcount over time; its website's traffic tier.company_news, company_headcounts, traffic_ranksDaily to every three days; headcount yearly; traffic monthly
Market seriesHow industries, countries, size bands and markets moved over 7, 30 and 180 days: hiring, seniority and work mode, pay, sales organisation, technology adoption and churn, go-to-market changes, announcements, intent surges, traffic tiers, headcount, listings; same-store cohorts, rates, growth and an index.market_seriesWeekly, as of each Sunday (section 9)

Every file carries company identifiers (company_id, company name, and ticker and exchange for listed companies), so the products join to each other.

2. Sources

All sources are public and first-party: the company publishes the information itself, or it is a public record.

  • Company websites. The homepage of each company's own domains, the tag-manager containers the homepage loads, and, when the homepage links to one, the careers page.
  • Public DNS records. Mail-server (MX) and TXT records of the company's domain. Companies publish these for email delivery and to verify ownership of their domain with software vendors.
  • Company job boards. The job boards companies publish so that job sites and search engines can pick up their postings, read through the board's public feed.
  • Company newsrooms and feeds. The newsroom or press page a company's homepage links to, and the RSS or Atom feeds the homepage declares: the company's own announcements, in its own words.
  • SEC EDGAR. Form D notices of exempt private offerings, filed by companies within 15 days of a first sale of securities; for listed companies, Form 8-K current reports (by item number) and the employee count stated in the annual report (Form 10-K).
  • Chrome UX Report (Google, CC BY 4.0), an open dataset: the monthly popularity rank bucket of each website's origin, used as a traffic tier. No visitor data of any kind.
  • Wikidata (CC0): founding year, headcount, headquarters, identifiers.
  • European Central Bank reference rates, to put salaries on one currency scale.

3. Collection principles

Collection is by an identified crawler, FokalsBot, which:

  • names itself in every request and links to a public page that explains what it reads and how to opt out;
  • reads and obeys each host's robots.txt, including crawl delays, before every request;
  • makes one request at a time to a host, with a pause between requests;
  • reads public pages only: it does not sign in, fill in forms or read anything behind a login;
  • never changes its identity. When a site's network protection refuses the direct read (an HTTP 401, 403 or 429, or a challenge page), the page may be fetched through a proxy address or a third-party fetching service, under the same name and the same robots.txt check; no challenge is solved by us, and a robots.txt "no" ends the read on every path. A site that refuses on every path is recorded as refusing and asked again a month later.

A small share of script-built homepages is read in a browser, within a fixed daily budget, to see which tags the page loads; the browser identifies itself the same way and records only the addresses of what the page loads.

SEC EDGAR is read within SEC's published fair-access rules, with the business name and a contact address in every request.

4. Coverage

Coverage is the Fokals company index: listed companies, the brands they own (each with its own website) and checked private companies. For each company the crawler reads the company's own domains.

  • Panel companies, and listed companies headquartered in a priority market, are read daily; other listed companies every three days; other companies, and brand sites without a known company, weekly (the latter are not delivered, see SOURCING.md). The priority markets are the United States, Canada, the United Kingdom, Ireland, the European Union, Switzerland, Norway, Australia, New Zealand and Japan. Work elsewhere runs at the lower cadence and after them; no country is excluded. A company's market is its headquarters country; failing that, the United States for an SEC registrant, the country of its ISIN, or the market its website was found in (the Chrome UX Report's country table it ranked in, or the country its domain ending names).
  • Only websites written in English are covered (section 4.2).
  • Some sites decline the direct read. They are read through a proxy address or a fetching service under the same rules (section 3) and less often; a site that declines on every path is left out, and the share that declines is reported in each release.
  • Job boards are found through each company's own website. Boards hosted on the major applicant tracking systems and job markup on companies' own careers pages are supported.

4.1 Listed companies

Every listed equity issuer worldwide is in the index, whether or not we found its website first. The list (listed_securities, one row per listing keyed by FIGI) is built from four open sources and refreshed monthly. OpenFIGI's public filter endpoint is read market by market, each market asked by the one exchange code that lists its issuers once (a country composite such as JP, or a single main venue such as LN; the United States through the NYSE American venue, whose rows cover every NMS stock, kept under their composite FIGI). Only an issuer's own equity counts: common and foreign shares, depositary receipts, REITs and partnership shares, never funds, exchange-traded products, rights, warrants or preferred stock. Each listing records its FIGI, share-class and composite FIGIs, ticker, venue code and MIC, and when it was first and last seen; a listing not seen for 45 days after a pass completes is marked delisted and kept, so history joins stay intact. Wikidata (CC0) supplies ISINs, LEIs, exchange-and-ticker pairs, websites and countries for the companies whose item we know, read by item id, the one path its robots.txt permits. The Global LEI Foundation's ISIN-to-LEI file and LEI records (CC0) supply LEIs, legal names and countries. The SEC's ticker file and company submissions supply the venue, CIK and website of US issuers. Listings sharing a share-class FIGI are one security; its primary listing is the home market's when the issuer's country is known from an ISIN or Wikidata, otherwise a market outside the United States, otherwise the composite record. Each security is tied to a company in this order: ISIN, LEI, ticker on the same market, then exact normalised name in the same country. An issuer with no match becomes a company (match_basis listed, confidence medium, status public) with its identifiers; identifiers on an existing company are filled only where empty, and the website is taken from Wikidata or, for US issuers, the SEC, never guessed. Every listing states how it was matched (isin, lei, ticker, name, wikidata) and which source asserted it. Listed companies' websites are read daily in the priority markets and every three days elsewhere (above). An ISIN or LEI is stored only when its check digit is valid.

4.2 Languages

Only websites written in English are read, labelled and delivered. The language of a website is that of its homepage as served to a reader who asks for English (Accept-Language: en). It is judged from the page's own words when there are enough of them: the script the page is written in, and the share of its words that are the commonest function words of English against those of twenty-three other languages. When the words do not decide (a page of names and numbers, a page built by scripts), the language the page declares (<html lang>) stands; country codes and three-letter codes written there are read as the languages they mean. A declaration never overrules words in another language: templates ship declaring English and are filled in German. The other way round the words must be clear, because English is on pages of every language (menus, cookie notices, browser warnings): a page that declares another language is English only when at least a tenth of its words are English function words, at least eight different ones, and no other language has a presence of its own. A page in two languages that declares the other one is that one's. Among closely related languages (Czech and Slovak, Danish and Norwegian, Croatian, Serbian and Bosnian) the declared one is taken. A page whose language cannot be told is read; exclusion needs evidence. The rules carry a version; when it changes, every website is judged again from the words stored with its last reading.

Checked by hand on 29 September 2026: of 61 excluded homepages read, drawn from the judgements most likely to be wrong, none was in English. Of the websites read as English, about one in a hundred was in another language under the first version of the rules; the second version, described above, was made to leave those out.

A site that publishes in several languages is read in its English version. When the homepage is in another language, the addresses the page itself gives for its English version are read, at most two, in this order of trust: the alternates declared in its head (hreflang), the links of its language switch that are marked with a language, and links named by the language ("English", or "EN" where the address says so too). Nothing is guessed: no address is tried that the page does not give. The version counts when it is on the same website (the site's own domain before another domain of the same name), answers, and is itself in English by the rules above; it is then the site's reading, every field comes from it, and later reads go to it directly. read_url names the address read and home_language the language of the homepage. Of the 88,850 websites whose homepage was in another language on 30 September 2026, 28,549 (32 percent) were read in the English version they declared; the share of websites left out fell from 38 to 26 percent.

A homepage in another language with no such version is not read further. Nothing of it is stored, labelled or delivered, its DNS records are not read, and it is asked again a month later. A candidate website of that kind is turned away on its pages alone, before any model is asked. A company with a website found in another language and none read in English is left out of every product, with its postings, signals, announcements, scores and sales figures, and of every cohort and count of the market series and the pay benchmarks; its job boards and feeds are not read while it is out. What was collected before stays on file and returns with the company when one of its websites is read in English. Readings made before the rule existed are judged the same way, from the words stored with them.

The list of listed securities (4.1) is not filtered by language: a listed company left out of the products keeps its listings and identifiers there, without a link to a company record. A site that offers English at an address its homepage does not link or declare is left out.

5. Marketing stack

5.1 Detection

Each reading combines up to five views of a site: the homepage itself, the tag-manager containers it loads, a browser read for script-built pages, the response headers, and the public DNS records. A technology is detected when its known signature appears in one of these views.

The catalogue holds 6,283 technologies in 68 categories. It is made of two lists, and every technology row says which one recognised it (technology_catalogue):

  • fokals: 209 technologies in 33 categories whose signatures Fokals writes and tests: advertising pixels and conversion tags, tag managers, web and product analytics, consent management, commerce and content platforms, marketing automation, email and SMS, CRM and account intelligence, chat and support, experimentation and personalisation, reviews, affiliate, payments and buy-now-pay-later, attribution, site search, loyalty, video, security and content delivery, workplace software and domain verifications. These carry public account ids, are tied to intent topics (section 7) and have weekly market series (section 9).
  • open: 6,074 technologies recognised by signatures from webappanalyzer, the open-source continuation of Wappalyzer, as published on 16 September 2026. It adds the long tail of business tools and the software a site is built on: JavaScript libraries and frameworks, web servers and programming languages, page builders, WordPress plugins and themes, Shopify apps, hosting, fonts, booking, recruitment and learning systems, and domain verifications with software and AI vendors.

From each entry of the open catalogue Fokals uses the signatures its readings can be held against: script addresses, the page as served, the scripts written into it, meta tags, single page elements, response headers and TXT and MX records. Only signatures the catalogue marks as certain are used. Signatures that need a visitor's browser session (script variables, cookies) are not used, and an entry known by nothing else is not in the count. An entry that is a product of the fokals list is left out, so no product is listed twice. Web standards and server settings are not counted as technologies.

A tag-manager container and the requests of a browser read show the tools a company set up and also what those tools load for themselves. From these two views an open technology is recorded only when it is a tool a company chooses (analytics, advertising, chat, consent and the like), not a library or a host behind one.

open technologies are in company_technologies and company_tech_events like any other. An open technology is established once 100 delivered companies run it, and it stays so. From then on its changes count for intent (section 7.1) and it has weekly market series (section 9). Below that line its numbers are a handful of websites: it is on the stack and in the change record only.

Where a technology carries a public account id (an advertising pixel id, an analytics property, a tag-manager container), the ids are recorded with it. seen_via says which views saw it.

5.2 Site facts

Each reading also records what the site offers: the markets and languages it declares for alternate versions of the page, the currencies of prices marked up on it, the company's own social accounts, its mobile apps, which key pages it links to (pricing, demo, free trial, contact sales, careers, investors, newsroom, partners, affiliate programme, store locator, subscription or membership, wholesale, gift cards, blog, developer documentation, status page) and the message in its announcement bar.

5.3 Website labels (sites-v2)

Each company website is labelled by the same evaluation model as the postings (section 6.3), from the page's own words (title, description, an excerpt of its text, its announcement bar), the key pages it links to, its markets, currencies and apps, the structured data on it and the technologies it runs. The labels are: industry (21 values), business model (b2b, b2c, both), offering (11), sales motion (4), price tier (5), target customer size (consumers, smb, mid_market, enterprise, mixed), pricing model (free, freemium, subscription, usage_based, one_time, quote_only, not_stated), geography scope (local, national, multi_country, global), primary persona (9), growth stage (startup, scaleup, established, incumbent, unknown); three 0–3 scores (e-commerce maturity, technology sophistication, content intensity), each the most probable of four described levels; and fifteen flags (sells online, physical locations, subscription, AI product, sustainability, enterprise focus, international, API, mobile app, free trial, demo request, compliance badges, partner programme, sells to government, regulated industry). A choice or score below 45% probability is left empty; a flag needs 70%. A site is labelled after its first full read, and again when what the model would read has changed and the labels are a month old, or the version changes.

5.4 Changes

  • The first reading of a site sets its baseline and writes no change. Every change in company_tech_events is one that was seen happen between two readings.
  • A technology or fact absent from a reading is first marked missing. It is recorded as removed only when it is still absent on a later reading at least 20 hours on. Tags come and go with consent banners and tests; this rule keeps those out of the change record.
  • Something can only be called missing by a reading that covered the view it was seen in: a DNS record is not removed because a DNS lookup failed, nor a tag because its container could not be read.
  • A commerce or content platform of the fokals list removed while another appears is recorded as a platform change.
  • The catalogue grows, and its signatures are revised. A technology is recorded as added only when a view it was found in had been read before with the signature that found it. The first reading of a website with a new or revised signature sets that technology's baseline there and writes no change; so does the first reading of a view that was not read before (the first browser read of a script-built page). first_seen_at is then the time of that reading. A technology that a revised signature no longer matches leaves the list without a removed event. No date in the change record is the date Fokals began to recognise something.
  • A new announcement-bar message is recorded as a promotion change; a bar that disappears is not.

6. Hiring

6.1 Postings

Every posting on a company's board is kept with its title, department, team, locations and country, work mode, employment type, posting date where the board gives one, and the dates we first and last saw it. From the posting text we read, by rule:

  • Pay: as stated in the board's structured fields or in the text (a pay range with a period or a pay keyword), and on one yearly US-dollar scale at the latest ECB reference rate. Yearly figures outside 5,000 to 2,000,000 US dollars are kept as written but not converted.
  • Minimum years of experience, degree required, visa sponsorship (offered or not offered), languages required.
  • Tools named: about 200 software products, from a catalogue matched on whole words, with common words matched only in context. A company's own products are not counted as tools it uses.

Contact details (email addresses, phone numbers, personal profile links) are removed from the text before it is stored. Posting descriptions are used to derive the fields above and are not delivered.

6.2 Opening and closing

  • A posting already on a board at our first read of that board is marked found_on_first_read: we did not see it open, so it is not counted as a new posting.
  • A posting that is no longer on a board is closed when it is still absent on a later complete read at least 20 hours on. A board that suddenly answers empty is read again before anything is closed.

6.3 Labels (jobs-v2)

Each posting is labelled by an evaluation model that answers typed questions with calibrated probabilities and generates no free text: job function (26 values), seniority (9 levels), remote policy (remote, hybrid, onsite, unspecified), contract type (permanent, contract, temporary, internship, unspecified), a 0–3 technical depth (the most probable of four described levels), ten flags (people manager, AI role, paid media, selects tools, international, new initiative, urgent, team build, replaces vendor, budget owner) and fifteen tool families the posting names or requires (Salesforce, HubSpot, Google Ads, Meta Ads, Google Analytics 4, Adobe, Shopify, SAP, Workday, Snowflake, Databricks, AWS, Azure, Google Cloud, Figma). A choice or score below 45% probability is left empty; a flag needs 70%. The labels, questions, model and floors are frozen under the version name and verified by fingerprint on every run; any change is published as a new version. jobs-v1 carried the function, seniority and the first six flags; every jobs-v1 key keeps its meaning in jobs-v2.

6.4 Daily hiring

company_hiring_daily is written once for each closed UTC day: open, new and closed postings, and the open postings by function, seniority, country and work mode, with AI roles, paid-media roles, tools named and median yearly pay.

6.5 Accuracy checks

Label accuracy is measured by hand. A person is shown a random labelled website or posting with exactly the evidence the model read (section 5.3, section 6.3) and grades every label field as correct, wrong or unsure, noting the right value where the model was wrong. Precision per field is correct / (correct + wrong); an "unsure" verdict is counted but decides nothing. The figure for a version is the same ratio over all its graded fields. Numbers are per label version (sites-v2, jobs-v2): a new version starts its own count, and a version's figure is quoted only once at least 100 websites or 100 postings under it have been graded. The graded rows are kept (label_checks) so a figure can be recomputed and audited.

6.6 Sales organisation (sales-v1)

A company's sales organisation is read from what it advertises. Two things happen to a posting in a sales function (the jobs-v2 functions sales and sales development, account-management titles filed under customer success, sales leaders filed under executives).

First, deterministic extraction from the posting's own text: on-target earnings and base pay where the posting names them apart (as written, and as yearly US-dollar midpoints at the reference rate), the base-to-variable split, whether commission is uncapped, the quota or target stated, the ramp period, the stated lead mix (inbound, outbound, mixed, and the inbound share when a percentage is given), the average deal size and the sales-cycle length. A figure is kept only within the plausibility bounds used for salaries. These are the employer's statements about the role: advertised pay, not pay received; a target, not attainment.

Second, one more read by the evaluation model under the frozen version sales-v1: the role (SDR or BDR, account executive, account manager or customer success, sales engineer, sales leadership, sales operations, partnerships or channel, other), the segment sold to (SMB, mid-market, enterprise, unspecified) and five flags (quota carrying, inbound leads, outbound prospecting, new business, founding sales hire). Floors, fingerprinting and versioning are as in section 6.3.

From these, company_sales_weekly is written once per company and closed week: open, new and closed sales postings, the sales share of all open postings, the open sales postings by role, segment and country, the median advertised base and OTE by role, the sales tools named, how many postings state a quota and the median stated inbound share; plus two dated events, a first sales posting in a country new to the company's sales hiring and a first enterprise-segment posting (both only for companies with earlier sales postings, so a board's first read does not count). sales_pay_benchmarks gives, per week, role and country, the quartiles of advertised base and OTE across companies, for groups of at least five postings. Sales postings also give three dated signals (sales_expansion_country, sales_upmarket, founding_sales_hire) that carry no intent topic under intent-v2: they are evidence in the signal feed, not a score.

7. Intent

7.1 Signals

A signal is a dated public action that indicates a coming purchase or build, tied to one or more of 69 topics in nine groups (advertising, marketing, commerce, bookings and events, sales and service, data and technology, website technology, business systems, growth). Signals come from:

  • Website changes: an advertising platform, tool or account added or removed, a platform change, a new domain verification with a software vendor, a new market, language, currency or app, a new sales motion (pricing, demo, free trial, contact sales), a new programme (affiliate, subscription, wholesale) or investor relations. A technology of the open list counts once it is established (section 5.1), under the topics of its category.
  • Job postings: a tool named in a posting (stronger when named in the title), and from the labels: paid-media roles, roles that select tools, international roles, AI roles, new initiatives, and leadership hires that build a function.
  • Funding: a Form D notice, weighted by the amount raised.
  • Sales organisation and leadership roles (section 6.6, section 8.2): a first sales posting in a country new to the company's sales hiring, a first enterprise-segment posting, a founding sales hire, and an executive or board change by role. These carry no topic under intent-v2: they appear in the signal feed and the datasets as dated evidence and score nothing.
  • Model-read topics (model_topic, weight 2, signal-v2): the evaluation model reads a posting that selects tools, builds a team, replaces a vendor or names software, and each labelled website, in two steps: first which of the nine topic groups (or none) the text shows the company adopting, buying, building or expanding, then, within that group only, which topics. A topic needs 75% probability; at most three are kept. One signal per posting, and one per website reading whose stable facts changed. The groups, questions, floors and model are frozen under the version name like the labels.

Each signal has a weight by its kind and strength. At most five signals of one kind are taken from one reading of a site, so a site that switches on dozens of markets at once does not dominate.

KindSourceWeight
ad_platform_addedsite2.5
tech_addedsite2 (a paid tool of the fokals list), 1 (any other, and any tool of the open list)
tech_removedsite, dns1.5 (fokals), 1 (open)
sales_motion_added, programme_added (affiliate), investor_relations_added, currency_addedsite1.5
dns_verification_added, market_added, app_addeddns, site2
replatformedsite3
tech_account_added, language_added, social_account_added, store_network_added, careers_added, programme_added (subscription, wholesale), developer_programme_addedsite1
posting_toolcareers2 (in title), 1
posting_paid_media, posting_selects_tools, posting_international, posting_new_initiativecareers1.5
posting_aicareers1
posting_leadershipcareers2
model_topiccareers, site2
fundingfiling1–4 by amount
sales_expansion_country, sales_upmarket (2), founding_sales_hire (3), leadership_role_change (2)careers, newsno topic

7.2 Weekly scores

For each company, topic and closed week, the score is computed from the signals of the 90 days to the week's end. Each signal's weight fades by half every 30 days; the faded weights are summed per topic and mapped to a 0–100 scale that rises quickly with the first strong signals and flattens as evidence accumulates. New funding raises the other topics of the same company. Scores below 5 are not written.

A surge is a score of at least 50 that is at least double the company's own average for that topic over the previous twelve weeks (or a first strong week). Each row carries its five strongest signals as evidence.

Weeks run Monday to Sunday, UTC. Each week is written once, after it closes, and never changed.

8. Funding, announcements and scale

8.1 Form D

Form D notices are read for the offering itself: issuer, industry group, state, amounts offered and sold, date of first sale, securities offered, exemptions claimed, number of investors, minimum investment and commissions. Pooled investment funds are left out of the product. The people named in a filing (executive officers, directors, promoters, sales recipients and signatories) are never read. A filing is tied to a company only when the issuer's normalised legal name matches exactly one company in the index; otherwise it is delivered without a company id.

8.2 Announcements (news-v1, sec-items-v1)

An announcement is something the company said about itself in public: an item on its newsroom or press page, an entry in an RSS or Atom feed its homepage declares, or, for a listed company, a Form 8-K current report. Each is stored once (company_news) with its title, its date as published, a short excerpt of the company's own text (at most 1,200 characters) and the address of the original.

Newsroom and feed items are read once by the evaluation model under the frozen version news-v1: one event type from a fixed list (product launch, partnership, acquisition made, acquired or merged, funding round, leadership change, layoffs or restructuring, expansion, incident, financial results, award or recognition, regulatory or legal, other) and five flags (about this company, executive appointment, executive departure, names an amount, mentions hiring), each with its probability; the choice is kept at 0.45 or above, a flag at 0.7 or above. An item the model does not read as being about the company itself carries no event type.

Form 8-K items need no model: the form's item numbers say what was filed, and each maps to an event type under sec-items-v1 (for example 1.01 material agreement → partnership, 2.01 → acquisition made, 2.05 exit costs → layoffs or restructuring, 3.02 unregistered equity sale → funding round, 5.01 → acquired or merged, 5.02 → leadership change, 1.05 → incident). The excerpt is the text under those items. An officer or director change is kept as the company's dated statement; no profile of the person is built or delivered.

For a leadership change, from either source, the model reads the company's text once more under the frozen version roles-v1 for the role concerned (chief executive, chief financial officer, chief operating officer, revenue or sales or commercial leader, marketing leader, technology or product leader, people leader, general counsel, board, other executive, none) and the direction (appointed, departed, succession), with a flag when more than one role changes. Only the role is stored (company_news.roles); the person is not. Each becomes a dated signal (leadership_role_change) with the role in its detail and no intent topic.

Every announcement with an event type is also a dated signal (news_event) with the topics the event implies (funding → new funding; expansion → international expansion; an acquisition or a leadership change → building a new function), weighted 1 to 3.

8.3 Headcount

Employee counts are kept over time (company_headcounts), one row per company, date and source: the count a listed company states in its annual report (Form 10-K, dated at the report's period end) and the count Wikidata records (dated as Wikidata dates it). The 10-K count is read from the company's own sentence ("we had approximately 12,345 full-time employees", "a workforce of about 3,000"): the first plausible count in a sentence about the company's own staff, ignoring sentences about customers' employees, per-employee figures and benefit plans. Each annual report is read once. A count is as the source stated it; no estimate is made between points.

8.4 Traffic tier

The Chrome UX Report publishes, for each website origin with enough Chrome visitors, the smallest popularity bucket it falls in each month: top 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, one, five, ten or fifty million origins. The tier of a company's website is the best bucket among its origin variants (with and without www, https and http), per month (traffic_ranks). Sites with too few visitors have no tier. The data is Google's, licensed CC BY 4.0, and is delivered with that attribution; it contains no visitor-level information.

9. Market series

The company-level tables are aggregated once a week into market series: how an industry, a country, a size band or a market moved over the last 7, 30 and 180 days (365 for headcount), as of the Sunday the week closed. A series is a metric, a dimension and a window; each row carries the count, the cohort it is over, a rate in the unit that fits the metric, the same count over the window just before, the growth between the two and an index against the series' first as-of date. Rows are written once per as-of date and never rewritten, so a buyer reading the table on any later day sees what we knew that Sunday.

Cohorts are same-store. The index of companies grows every day, so a raw count would mostly measure our own coverage. A rate is therefore taken over companies (or websites) covered since before the window began and still covered at its end, and the previous window's count is taken over the same cohort. A website's first reading writes no change events, so adoption counts carry no discovery noise either. Technology added, removed and prevalence are written for the 209 technologies of the fokals list and for each open technology once it is established (section 5.1). An open technology's series begin with the first window that lies wholly after it was established and at least 14 days after it was first counted: in those days each website is being read with its signature for the first time, and a count across them would rise with the reading, not with adoption. For the tools recognised only by a public DNS record (a workspace, an email provider, a domain verification) the series begin with the first as-of written after 3 October 2026, when their changes were first counted; no series written before that date changed.

A row over fewer than 20 companies (20 postings for pay and share metrics) is not written: a small cohort says more about our coverage than about the market.

Units. Events a company does once (adopted a technology, raised money, cut staff, opened a country, added a pricing page) are stated per 100 companies or per 100 websites. Things counted in postings (AI roles, remote roles, seniority, function, sales roles, segments) are stated as a share of the postings opened in the window. Money is a median in yearly US dollars. Headcount change is the median year-over-year percentage over companies with two counts at least 300 days apart. Listings are plain counts with no cohort; a new listing is one that appeared, in OpenFIGI's monthly lists, on a market whose list was already held for a week or more, so the first load of a market is never counted as new. Growth is (count − previous) / previous and is empty when there was nothing before; the index is 100 at the series' first as-of date and the rate relative to it after.

Resolution follows the source. Hiring, technology, sales, announcements, intent and website changes are daily, so every window is real. Traffic tiers are monthly: the 30-day window compares the latest month with the one before, the 180-day window with the month six earlier. Headcount is yearly. History starts on 27 September 2026 and cannot be backfilled: the 30-day windows are complete from late October 2026, the 180-day windows from late March 2027.

Dimensions. Industry is the website label (sites-v2, 21 industries); sector is the company's Wikidata or DBpedia sector; country is the headquarters country; size is the band of stated employees (1–10, 11–50, 51–200, 201–1,000, 1,001–5,000, 5,000+); market is the exchange code of a listing. Every series also exists over all companies. The metric catalogue (family, unit, windows, dimensions, cohort) is published by the API and in DATA_DICTIONARY.md; a change to a metric's definition is a new metric id, never a rewrite.

10. Time conventions and history

All times are UTC. Daily and weekly rows are written once after the period closes and never rewritten; a period written more than seven days late is flagged reconstructed. Website changes, postings and signals carry the time they were observed.

One exception is on record. The market series as of 27 September 2026, the first as-of and the base of every index, and the sales pay benchmarks of the week of 21 September 2026 were written hours before the language rule (section 4.2) took effect. Both were computed again, twice in the same week, before any delivery and before a later period was written: once when the rule left the websites in other languages out, and once more when the reading of the English versions those websites declare brought a third of them back. The first period and every later one are thereby over the same population. Each rewrite is recorded with the row counts before and after.

A second correction to the same as-of: its listings_new rows counted the first load of the listed securities, on 26 September 2026, as 72,776 new listings. Those rows were removed on 4 October 2026, before any delivery, and the count was defined as it stands in section 9. The listings_new series therefore begins with the as-of of 4 October 2026.

11. Known limitations

  • Tags loaded only after a visitor accepts cookies, or only on pages other than the homepage, may not be seen.
  • Sites that decline the direct read are read less often; sites that decline on every path are not covered.
  • Websites with no English homepage and no English version they declare, and companies with no such website, are not covered, whatever their market. Coverage of markets where English is not the first language is therefore partial, and skews to companies that publish in English. A site read in its English version is described by that version: a tag or a page that exists only on the version in the home language is not seen. Language detection has an error rate: a short or mixed-language page can be judged wrongly either way.
  • A company whose job board is not linked from its website, or is hosted on an unsupported system, has no hiring data.
  • A company whose homepage links to no newsroom and declares no feed has no announcements unless it is a listed company (Form 8-K). A newsroom page without dates gives items dated when first seen. Event types on newsroom items are model output with a known error rate; 8-K event types follow the form's item numbers and are only as specific as those.
  • Headcount has a point only where a public statement gives one: annual reports of listed companies and Wikidata. Traffic tiers exist only for origins with enough Chrome visitors, and a bucket is a range, not a visitor count.
  • Tools named in postings are matched from a catalogue; unlisted tools are not counted.
  • Signatures of the open list are kept in the open by its contributors and are not tested one by one by Fokals. They are less exact than those of the fokals list: a library can be named for a file that only shares its name, and a product for a link to its maker. technology_catalogue lets a reader keep to either list.
  • A technology recognised only by what a visitor's browser holds (a script variable, a cookie) is not seen.
  • Sales facts (pay, quota, ramp, lead mix) are what the employer advertises in the posting, not what representatives earn or attain; a posting that states nothing has no facts. Roles and segments are model output with a known error rate.
  • Labels are model output with a known error rate; accuracy figures are published after the first hand check.
  • Intent scores are our assessment of public signals, not statements about a company's plans, and not forecasts.
  • Funding matching is conservative: filings under a legal name that differs from the company's known names are not tied to it.

12. Change management

Label and intent versions are frozen by name. A change to labels, questions, floors, topics, weights or the scoring formula is released under a new version, with at least 90 days' notice for breaking changes. Additions to the technology and software catalogues are listed in the changelog (section 14).

13. Versions

VersionSinceWhat it covers
jobs-v12026-09-22Posting labels: function, seniority, six flags.
jobs-v22026-09-23jobs-v1 plus remote policy, contract type, technical depth, four flags, fifteen tool families.
sites-v12026-09-22Website labels: industry, business model, offering, sales motion, price tier, seven flags.
sites-v22026-09-23sites-v1 plus customer size, pricing model, geography, persona, stage, three scores, eight flags.
signal-v12026-09-23Model-read buying topics (model_topic signals).
signal-v22026-10-05signal-v1 over the nine groups and 69 topics of intent-v2.
brief-v12026-09-23AI-written company notes, cited to facts (not delivered in the data products).
intent-v12026-09-22Weekly intent scoring.
intent-v22026-10-05intent-v1 with 17 more topics and two more groups, and signals from established open technologies. The 52 topics of intent-v1 keep their ids, names and groups; the scoring is unchanged.
news-v12026-09-24Announcement event type (thirteen kinds) and five flags, read from the company's own text.
sec-items-v12026-09-24Form 8-K item numbers mapped to the same event types (no model).
sales-v12026-09-25Sales postings: role (eight), segment (four), five flags.
roles-v12026-09-25Leadership changes: the role concerned (eleven) and the direction (three); never the person.

The technology catalogue is at version 2 (section 14).

A website is labelled under the current version on its next read when its stored version differs. Postings keep the version they were labelled under until they close.

14. Changelog

Additions to the technology and software catalogues, and changes that need no new label version.

DateChange
2026-10-05Technology catalogue version 2. The open list is added: 6,074 technologies recognised by signatures from webappanalyzer (commit eea872a of 16 September 2026), beside the 209 of the fokals list (version 1). technology_catalogue is added to company_technologies and company_tech_events, and catalogue to the technology objects of the API. On each website the first reading with the new entries is their baseline and writes no change (section 5.4). The company summary's technologies and tech_count count both lists.
2026-10-05intent-v2 and signal-v2. 17 topics and two groups (bookings and events, website technology) are added to the 52 topics and seven groups of intent-v1, which keep their ids, names and groups; the scoring is unchanged. Each open technology is tied to the topics of its category. An open technology is established once 100 delivered companies run it: from then on its changes are signals (weight 1, or 2 for a domain record) and it has weekly series, which begin with the first window wholly after it was established and at least 14 days after it was first counted. The two latest closed weeks are also written under intent-v2, from the same signals, so the scores continue without a gap. scored and established_at are added to /api/v1/technologies.