Compliance

Data sourcing and compliance

We collect only what companies publish themselves, from public sources, with a crawler that identifies itself and obeys robots.txt on every request.

Principles

How we source data

.01

Public sources only

We collect public pages and public records. Nothing is collected from behind a login, a form or a paywall.

.02

First-party by design

Every record comes from the company itself or from a public filing. We do not buy or collect data from brokers, platforms, marketplaces or social networks.

.03

No personal data

The datasets describe companies, not people. Contact details are removed at collection, and individuals named in filings are not read.

.04

No MNPI

Every input is public at the time it is observed. We receive nothing from insiders, from companies under confidentiality or from any non-public source.

Collected

  • Company homepages, the tag-manager containers they load, and public DNS records
  • Job postings on the boards companies publish for distribution
  • Announcements on company newsrooms and feeds
  • SEC filings: Form D, Form 8-K and the employee count in annual reports
  • Open reference data: OpenFIGI, GLEIF, Wikidata and the Chrome UX Report

Never collected

  • Anything behind a login, a form or a paywall
  • Personal contact details, profiles of individuals or social network data
  • Visitor, cookie, device or other consumer data
  • Employee-reported or crowdsourced data
  • Data from brokers, platforms or marketplaces

How FokalsBot behaves

Collection rules are enforced in code, in a single component that every request passes through.

Identified
Every request names FokalsBot and links to a public page describing what it reads and how to stop it. It never changes its identity.
robots.txt
Read and obeyed before every request, on every path, including crawl delays. An opt-out takes effect within a day.
Rate
One request at a time per host, with pauses between requests. Most websites are read once a day to once a week.
Public records
SEC EDGAR is read within its fair-access rules, with the business and a contact address named in every request.
Requests

For site owners and companies

Stop the crawler
Add a User-agent: FokalsBot rule to your robots.txt. Instructions are on the FokalsBot page.
Request removal
Write to us through the contact form naming the company. We review every request and confirm the outcome by email.
Privacy
How we handle enquiries and account details is set out in the privacy notice.

Due diligence questions

Is any input material non-public information?

No. All inputs are public at the time they are observed. We receive no information from insiders, from companies under confidentiality or from any non-public source.

Do you collect or deliver personal data?

No personal data is delivered. Contact details in job descriptions are removed before storage, and the individuals named in SEC filings are not read. A leadership change is recorded as the company’s dated statement and the role concerned.

How do you handle websites that restrict automated access?

robots.txt governs every request on every path. Where a site’s network protection refuses a direct request, the page may be retrieved through a proxy or a third-party fetching service, under the same crawler identity and the same robots.txt rules. We do not solve challenges, and a site that refuses on every path is not collected.

Where do intent scores come from?

From the public actions of companies: website changes, job postings, funding filings and their own announcements. They are not derived from browsing behaviour, bidstream data, cookies or any consumer data.

How are machine-learning models used?

An evaluation model answers typed questions about a page or a posting under a named, frozen version that is verified on every run. Accuracy is measured by hand for each version. Model output is identified as such in the data.

What documentation is available for a compliance review?

The sourcing and compliance statement, the methodology and the data dictionary are public. We complete due diligence questionnaires on request.