A team that needs data on many companies can build web crawlers and collect it, or license a feed that someone else collects. Both are sound choices. This comparison sets out what each takes in five areas: engineering, maintenance, conduct, structure and labelling. Scrapy and Playwright stand for the build, as two open-source tools a team could choose, and Fokals stands for the licensed feed. A build is the tool for a handful of known sites and for fields specific to your own question. A licensed feed delivers the same structured fields for many companies every day, with entity resolution, dated changes and versioned labels already in place. By the end you can list the work a build would take for your own sites and decide which parts to build and which to license.
What Scrapy and Playwright give you
Scrapy describes itself as an application framework for crawling websites and extracting structured data. It is an open-source Python framework, and its architecture has the parts a crawl needs: a scheduler that queues requests, a downloader that fetches pages, spiders that you write to parse the responses, and item pipelines that clean, validate and store what the spiders extract. Requests are scheduled and processed asynchronously, and feed exports write JSON, CSV or XML.
Playwright describes itself as web automation for testing, scripting and AI agents: one API that drives Chromium, Firefox and WebKit, available for TypeScript, Python, .NET and Java. Its repository shows the Apache-2.0 licence. A crawler uses it for pages that build their content with scripts, where the data is not in the page as first downloaded. Its network API lets a script watch every request a page makes, which is how you see the tags a page loads. Scrapy's guide to dynamic content recommends the scrapy-playwright plugin where a headless browser is needed.
Between them the two tools cover fetching, rendering and the plumbing of a crawl. What a page means, whether it has changed, and which company it belongs to are yours to work out.
Building and licensing side by side
| Area | Building with Scrapy and Playwright | Licensing from Fokals |
|---|---|---|
| Engineering | Spiders for each kind of source, a browser path for script-built pages, queues, storage and monitoring | A loader for a REST API and bulk files in JSON, JSON Lines or CSV |
| Maintenance | Layouts, job boards and blocks change, and each break is yours to notice and repair | A documented schema under named, frozen versions, with at least 90 days' notice of a breaking change |
| Conduct | A policy you write and enforce: identity, robots.txt, pace, what is never read | Collected from first-party company sources and public records, processed in-house, under a published sourcing statement |
| Structure | Company matching, baselines, change detection and closing rules to design | Datasets keyed on one stable company ID, every record with the time it was observed, every change a dated event |
| Labelling | A classifier to choose, version and check by hand | Job function, seniority and role flags on every posting, under named, frozen versions |
| Scope | The sites, pages, fields and cadence you choose and staff | Listed companies worldwide, the brands they own and verified private companies, refreshed daily |
| Cost | Engineers' time, at the start and every week after. The tools are open source | A licence by written agreement for internal use, embedding in a product or redistribution |
Engineering: kinds of source, not pages
The unit of work is a kind of source, not a page. Company data that covers technology, hiring, announcements and funding draws on at least four: company websites, careers pages and job boards, newsrooms and feeds, and regulatory filings. Each has its own discovery problem. A job board must be found before it can be read, and boards run on different applicant tracking systems, each with its own format. A licensed feed has that work done. Fokals delivers every role a company publishes in one Job Postings schema, labelled by job function, seniority and ten role flags, with location, work mode and advertised pay on one annual US-dollar scale.
Scale adds a second layer. Scrapy's common practices page says the framework has no built-in facility for running a crawl across several servers, so distribution is a design decision of yours. Scrapy's guide to dynamic content suggests a headless browser where crawling speed is not a major concern, so plan the browser share of a crawl on its own. With a licence that capacity is the vendor's to run: Fokals data arrives refreshed daily, and your side of the work is a loader for an API and files.
Conduct: robots.txt, RFC 9309 and a policy of your own
Conduct is the part no framework decides for you, and RFC 9309, the Robots Exclusion Protocol published in 2022, sets the floor. A crawler finds the group in robots.txt that matches its product token and obeys it. If the file cannot be reached because of a server or network error, the crawler must assume that everything is disallowed. A crawler may cache the file but should not use a cached copy for more than 24 hours, unless the file is unreachable. The RFC also states that its rules are not a form of access authorisation, so obeying the file is good conduct and settles nothing about what you may do with what you read. None of this is legal advice.
Scrapy has the switch. With ROBOTSTXT_OBEY enabled, a downloader middleware filters out the requests robots.txt forbids. Its settings reference lists that setting as True, with a one-second download delay and one concurrent request per domain, in the project that scrapy startproject generates. The fallback values built into the framework are False, no delay and eight, so read what your own project sets. A policy written down looks like this, for an illustrative crawler run by Acme Robotics:
# settings.py
USER_AGENT = "AcmeRoboticsBot/1.0 (+https://www.example.com/bot)"
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1
AUTOTHROTTLE_ENABLED = True # lengthen the delay when the site answers slowlyWith a browser, plan to make the robots.txt check yourself before each navigation, and to apply your own pause and user agent. Playwright's documentation shows how to override the user agent and how to abort requests by route.
The harder decisions come when a site blocks you. Scrapy's common practices page describes two courses: making yourself known through USER_AGENT so that a site's owners can reach you, and techniques that make traffic resemble a regular visitor's, such as rotating user agents and spreading requests over a pool of IP addresses. Which of them your crawler may use is a matter for your policy, and a due diligence review will ask.
A licence turns that review into a reading of documents. Fokals data is collected from first-party company sources and public records and processed in-house: what companies publish on their own websites and careers pages, what they announce and what they file. It is company-level data throughout. The sourcing statement and the compliance page are the documents to hand to a reviewer, and they hold the detail a due diligence questionnaire asks for.
Structure and labelling: from pages to rows
A fetched page is not yet a fact. Three decisions turn pages into data that can be tested. A build has to make each of them, and the Fokals methodology documents how the licensed feed settles them.
- Baselines. The first time a crawler sees a site or a job board, everything on it looks new. In Fokals data a first observation sets a baseline and is never counted as a change, so an adoption or a new posting in the record is one that happened under observation.
- Absence. A tag missing from one fetch may be sitting behind a consent banner, so a build needs a rule for when absence becomes a removal. In Fokals data a removal is confirmed before it is written, which keeps tags that come and go with consent banners and tests out of the record.
- Time. A series can be tested only if yesterday's rows stay as they were written. Fokals daily and weekly datasets are written once, after the period closes, and are never revised, and every record carries the time it was observed. The record is point-in-time by construction.
Identity comes next. A page belongs to a domain, a domain to a company, and a listed company to a security. Fokals keys every dataset on one stable company ID and carries the ticker, MIC, ISIN, LEI and share-class FIGI of a listed company, with its SEC CIK where it has one, and a brand or subsidiary carries the identifiers of its listed parent. In a build this is an entity resolution pipeline of its own, with check digits to validate.
Labels come last: counting engineering roles or software companies needs a classifier, thresholds, a version and a check. Fokals labels every posting by job function, seniority and ten role flags, and classifies company announcements and regulatory disclosures into 13 event types, each label produced under a named, frozen version. If you build, each of those choices is yours to make and to hold steady, because a label that drifts breaks every series built on it.
Maintenance: the cost that estimates leave out
Maintenance begins the day after a crawler first runs. A layout changes, a company moves its job board to another system, a site starts refusing your requests, a feed goes quiet. Each looks like a change in the data until someone finds the cause.
In a build the answer is alerts, runbooks and a share of someone's week, every week. Under a licence that upkeep is the vendor's, and what reaches you is a documented schema that holds steady. Fokals produces its labels and scores under named, frozen versions and gives at least 90 days' notice of a breaking change, so a pipeline or a product built on the feed keeps running while the web underneath it changes.
What a build is for
A build is the tool for three kinds of work.
- A few known sites. A few dozen named pages are a small Scrapy project, and you control exactly what is read and when.
- Fields specific to your question. The wording of one pricing page, or a page that matters to your thesis alone, is a field you define, so you collect it yourself.
- Collection as the product. If reading the web is what your company sells, the crawler is the asset.
What a licence is for, and how the two combine
A licence is built for breadth and for evidence. It fits when you need the same fields for many companies every day, when identifiers and labels matter more than raw text, when the data feeds a back-test or a product and must stay as it was written, and when a reviewer will ask how it was collected. Fokals answers each of those in its specification: firmographic, technographic, hiring and intent data on one company index, every observation dated, daily and weekly datasets written once and never revised, and sourcing documented in public.
The two combine well. License the base, meaning identifiers, postings, technologies and events, and point your own spiders at the few pages that are specific to your question. Fokals is delivered direct, by REST API and as bulk files in JSON, JSON Lines or CSV, which you load beside what your crawlers collect and join on the company's domain. The guide to alternatives to building your own crawlers covers the other options.
Frequently asked questions
Should you build a web crawler or buy company data?
Build when the sites are few and known, or when the fields are specific to your own question. License when you need the same structured fields for many companies every day, with identifiers, dated changes and versioned labels, and when you would sooner hand a reviewer a published sourcing statement than write and defend a crawling policy. Many teams license a base feed and build crawlers only for the pages that are theirs alone.
What is the difference between Scrapy and Playwright?
Scrapy is a Python framework for crawling: it schedules requests, downloads pages, runs the spiders you write and passes what they extract through pipelines. Playwright automates browsers, Chromium, Firefox and WebKit, through one API. A crawler uses Scrapy for the bulk of its fetching and a browser for pages that build their content with scripts. Scrapy's documentation recommends the scrapy-playwright plugin for combining them.
Does Scrapy obey robots.txt?
It does when ROBOTSTXT_OBEY is enabled, in which case a middleware filters out the requests that robots.txt forbids. Scrapy's settings reference lists the setting as True in the project that scrapy startproject generates and False as the fallback built into the framework, so a project created another way may have it off. Check your own settings, and set a user agent that names your crawler.
What is RFC 9309?
RFC 9309 is the IETF's specification of the Robots Exclusion Protocol, published in September 2022. It defines how a crawler reads robots.txt: which group of rules applies to it, how Allow and Disallow lines are matched, what to do when the file is missing or cannot be reached, and how long a copy may be cached. It states that the rules are not a form of access authorisation.
What does licensed company data include that a crawler does not?
Structure and upkeep. A crawler returns pages. A licensed feed such as Fokals returns rows keyed on a stable company ID, with observation times, changes measured against a baseline, labels under named, frozen versions and the ticker, ISIN, LEI and FIGI of listed companies, together with a published sourcing statement. The engineering that produces those properties, and the upkeep that preserves them every day, come with the licence.
Can licensed company data be used for back-testing or inside a product?
Yes, when the record is point-in-time and the licence covers the use. Fokals daily and weekly datasets are written once, after the period closes, and are never revised, and every record carries the time it was observed, so a test reads the data exactly as it stood on any day it covers. Labels and scores hold steady under named, frozen versions. Licences are written for internal use, for embedding in a product or for redistribution.
The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.
What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.