1. Summary
The company-signals products (Marketing stack, Hiring, Intent, Announcements and scale, and the Market series aggregated from them) are built only from public, first-party information: what companies publish on their own websites, in their public DNS records, on their newsroom pages and feeds and on the job boards they publish for distribution; US public records filed with the SEC (Form D, Form 8-K, annual reports); and two open datasets, Wikidata (CC0) for company facts and Google's Chrome UX Report (CC BY 4.0) for a website's monthly traffic tier. No data is bought from, or collected from, any platform, social network, marketplace or data broker, and no personal data is delivered.
2. How we collect
- Identified. Every request names our crawler, FokalsBot, and links to a public page that explains what it reads and how to opt out.
- Permission-respecting. robots.txt is read and obeyed before every request, including crawl delays. A site's opt-out takes effect within a day.
- Polite. One request at a time per host, with pauses between requests; most sites are read once a day to once a week.
- Public only. No sign-in, no forms, nothing behind a login or paywall.
- English only. A website whose homepage is not in English is read in the English version the page itself declares or links, on the same rules as every read, at most two addresses. Without one it is not read beyond the homepage, and nothing of the page is kept; it is asked again a month later. How the language is judged is in
METHODOLOGY.mdsection 4.2. - Blocked reads. When a site's network protection refuses the direct read (an HTTP 401, 403 or 429, or a challenge page), the page may be fetched through a proxy address or a third-party fetching service. The crawler's identity does not change, robots.txt is read and obeyed on that path as on every other, and no challenge is solved by us. A site that refuses on every path is recorded as refusing and asked again a month later. Sites read this way are read less often.
- Discovery. Candidate websites come from links on pages we read, the most-visited websites in the Chrome UX Report's monthly origin lists, worldwide and per country (Google, CC BY 4.0), the Common Crawl index (open web data), the job boards companies publish, and open company lists (DBpedia, the SEC ticker file, Wikidata, OpenFIGI's listed securities, GLEIF's legal entity identifiers). Each candidate is read once, on the rules above, and judged by the evaluation model before it becomes a company in the index.
- Announcements. A company's newsroom or press page (when its homepage links to one) and the RSS or Atom feeds its homepage declares are read on the rules above, at most three per company, daily to every three days. Each announcement is kept as a title, a date, the address of the original and an excerpt of at most 1,200 characters of the company's own text; the event type is read by the evaluation model under a frozen version.
- Public records. SEC EDGAR is read within SEC's fair-access rules, identifying the business and a contact address in every request: Form D, and for listed companies Form 8-K current reports (by item number) and the employee count stated in the annual report.
- Open data. Wikidata (CC0) for company facts; the Chrome UX Report (Google, CC BY 4.0) for the monthly popularity rank bucket of each website, delivered with attribution.
3. What we do not collect or deliver
- Personal data. Contact details in job descriptions (email addresses, phone numbers, personal profile links) are removed before storage. The individuals named in SEC filings (officers, directors, promoters, sales recipients, signatories) are never read.
- Small sites that may belong to individuals. Websites, postings and signals are delivered only for companies in our index; a brand's site that we cannot tie to a company is read for internal quality work only and never delivered.
- Employers' text. Job descriptions are used to derive structured fields and are not delivered.
- Websites in other languages. A website with no English homepage and no English version is neither stored nor delivered, and a company whose websites are all of that kind is left out of every product. The list of listed securities is not filtered by language.
- Visitor data. We read what a site publishes, not who visits it. The products contain no cookie, device or visitor-level data of any kind; the traffic tier is a monthly rank bucket of the website as a whole, published by Google under an open licence.
- People in announcements. A leadership change (a press release, or Form 8-K item 5.02) is kept as the company's dated statement with the excerpt the company published. The role concerned (chief financial officer, head of sales, board) is labelled; the person is not stored. No profile of a person is built, no person is tracked from one company to another, and no social network is read.
- Sales representatives. What we publish about a sales organisation is what the employer advertises: roles, segments, pay ranges, quotas, ramp and lead mix as stated in its own postings. We collect no ratings, reviews or reports from employees, and no attainment or earnings figures.
4. Intent data
Intent scores are derived only from the public actions of companies (website changes, job postings, funding filings, the company's own announcements). They are not built from the browsing behaviour of people, bidstream data, cookies or any consumer data.
5. Material non-public information
All inputs are public at the time they are observed. We do not receive information from insiders, from companies under confidentiality, or from any non-public source.
6. Opt-out and removal
A site owner can stop FokalsBot with robots.txt (User-agent: FokalsBot / Disallow: /). A company can ask for its information to be removed by writing to the contact address on the crawler page.
7. Governance
Collection rules are enforced in code, in one component that every request passes through. Label and scoring versions are frozen and verified on every run. Daily and weekly outputs are written once and never rewritten. AI-written company briefs are marked as such and carry the facts they cite; they are our assessment of public facts, not statements about any company's plans. The documentation is public: this statement, the methodology, the data dictionary, the API reference and the crawler page are published on the Fokals website.