Alternative

Alternatives to building your own company data crawlers

Four routes to rows of company data from the web: your own crawler, a managed scraping service, open datasets or a licensed feed. What each takes off your hands, and what stays with you.

Updated 5 October 20269 min read

This guide is for the engineering lead who runs web crawlers for company data, or is about to build them, and has to decide how much of that work to keep. It compares four routes to the same rows: an in-house crawler on Scrapy and Playwright, a managed scraping service, open datasets and a licensed feed. For each it says what the route takes off your hands and what stays with you.

The decision turns on what your team wants to own. A crawler of your own is built for the rows nobody sells: any site, any page, any field, on your schedule, with a method that belongs to you. A licensed feed is built for the team that wants the data and not the crawl. Fokals delivers firmographic, technographic, hiring and intent data on public and private companies worldwide, refreshed daily, with the change rules, the entity resolution and the labels already applied, so your engineers start from finished records.

What the in-house stack gives you

Scrapy describes itself as a framework for crawling websites and pulling structured data out of their pages. It is open source, written for Python, released under the BSD licence and, according to its site, maintained by Zyte with more than 500 other contributors. Its documentation covers spiders, item pipelines and feed exports to storage including local files, FTP, Amazon S3 and Google Cloud Storage. The settings reference documents ROBOTSTXT_OBEY, which makes Scrapy respect robots.txt policies when enabled, and DOWNLOAD_DELAY, the minimum wait between two requests to the same domain.

Playwright controls Chromium, Firefox and WebKit through a single API, headless or headed, on Linux, macOS and Windows. It is a Microsoft project, and its repository carries the Apache-2.0 licence. The library documentation describes APIs for launching and interacting with browsers outside a test runner. The two tools meet where a page builds its content with scripts: Scrapy's guide to dynamically loaded content points to a headless browser and recommends the scrapy-playwright integration.

Both are tools for fetching and parsing. What counts as a change, when a posting has closed, which company a website belongs to and how a row is labelled are rules you write on top.

The work beyond fetching

Requests are the small part of a crawler for company data. The list below is the rest, with what a licensed feed delivers in each case, as the Fokals methodology documents it. Use it as a checklist for a build or as questions for a vendor.

  1. Baselines. The first reading of a site or a careers page holds everything already there. In Fokals data a first observation sets a baseline and is never counted as a change, so every event in the record was observed to happen.
  2. Missing is not removed. A tag can vanish for a day behind a consent banner or an experiment. A removal is confirmed before it is written, so a technology that comes and goes stays out of the record of changes.
  3. Closing a posting. A role that leaves a careers page may be back the next day. A closure is confirmed before it is written, and the daily counts of open, new and closed postings are written once, after the day closes.
  4. Detection across the whole surface. A technology can show in page content, in scripts, in network requests, in response headers or in DNS records. Detection covers all five, and each technology carries its first-seen and last-seen dates.
  5. Entity resolution. Websites are tied to companies and brands to their listed parents on one stable company ID. A listed company carries its ticker, MIC, ISIN, LEI and share-class FIGI.
  6. Labels. Job function, seniority and role flags are produced under named, frozen versions, and a breaking change ships as a new version with at least 90 days' notice.
  7. The record. Every observation is dated. Daily and weekly datasets are written once and never revised, which makes the record point-in-time by construction.

Conduct is part of the build

Whoever sends the requests answers for them. The Robots Exclusion Protocol is specified in RFC 9309, published in 2022, which sets out rules that crawlers are requested to honour and states that those rules are not a form of access authorisation. A robots.txt file says what a site owner asks of crawlers, and honouring it is the crawler's part. What your crawler does when a site refuses it, whether it signs in anywhere and how it names itself are decisions for you and your counsel.

A licensed feed moves that work to the vendor, and the vendor's documents carry it into your review. Fokals publishes a sourcing statement and a compliance page written for a buyer's due diligence. A team that builds should be able to write the same pages about its own crawler.

Four routes side by side

The table repeats what each project or service states on the pages linked here. The Fokals row follows the data dictionary.

OptionWhat you getSourcesDeliveryLicence or terms
Scrapy, in-houseAn open-source crawling and scraping framework for PythonThe sites your spiders visitFeed exports in formats including JSON, JSON lines, CSV and XMLBSD licence
Playwright, in-houseA browser automation library for Chromium, Firefox and WebKitThe pages your script loadsWhatever your script writesApache-2.0 licence
Zyte API, managedA web scraping API that avoids bans and enables browser automation and automatic extractionThe sites you requestAPI responsesA paid service, priced per successful response
Apify, managedA platform that runs Actors, serverless programs for scraping and automationThe sites an Actor visitsDatasets for export; runs by API or on a scheduleFree and paid plans; usage billed in compute units
Common Crawl, open dataAn open archive of crawled web pages, free to useIts own crawl of the webWARC, WAT and WET files on Amazon S3, free to download over HTTPSIts terms of use; crawled content may carry its owners' terms
Wikidata, open dataA free, collaborative, multilingual knowledge baseEditors and bots, with sources recordedA query service and APIsCreative Commons CC0
Fokals, licensed feedFirmographic, technographic, hiring and intent data on one company index, refreshed dailyFirst-party company sources and public records, processed in-houseREST API and bulk exports as JSON, JSON Lines or CSV, delivered directWritten agreement for internal use, embedding in a product or redistribution

Managed scraping services

A managed service takes over the fetching. Zyte API is described in its documentation as a web scraping API that avoids bans, enables browser automation and enables automatic extraction, and its pricing page charges by successful response. Apify runs Actors, which its documentation defines as serverless programs in its cloud: each takes JSON input and carries out a job such as scraping a site or driving a browser. Each run keeps its results in a dataset that can be exported. Its pricing page lists free and paid plans.

What this route takes over, as those pages describe it, is avoiding bans, automating a browser and running the scraper in a cloud. What stays is the list above: a fetching service hands back pages or extracted fields, and turning them into dated changes tied to companies is still your code. The conduct questions stay too, since the requests are made on your instruction.

Open datasets

Some rows need no crawl, because someone already publishes them.

  • Web pages. Common Crawl, a non-profit, publishes its crawls of the web as an open archive that is free for anyone to use. Its get started page describes WARC, WAT and WET files hosted on Amazon S3, and its terms of use say that crawled content may be subject to separate terms from the owners of that content.
  • Company facts. Wikidata is a collaborative knowledge base whose data is published under the Creative Commons CC0 dedication.

Daily monitoring of named pages is a different rhythm: Common Crawl's site states that it adds 3 to 5 billion new pages each month, which is not the same as reading one company's careers page every day. Open data supplies reference facts and research corpora. A dated record of what each company does, day by day, is the work of a crawler or a feed.

A licensed feed

A feed takes over the whole list: fetching, change rules, entity resolution, labels and the record itself. You receive finished records on a documented schema, with the documents that say how the data is made, and your engineers spend their time on what the data is used for. Sample data for the companies you track is sent on request with the data dictionary and methodology, so the evaluation starts from the companies you track.

In Fokals the output of a typical company crawler maps to these datasets.

What your crawler producesFokals datasetWhat it delivers
Technologies detected on a siteTechnology StackEach technology with its category and its first-seen and last-seen dates
A log of site changesTechnology ChangesA dated event for every adoption, removal and platform migration, and for every new market, language, currency or app
Facts about each websiteWebsite ProfileThe markets, languages, currencies, apps and key pages of each website
Postings read from careers pagesJob PostingsEvery role a company publishes, labelled by job function, seniority and ten role flags, with location, work mode, advertised pay and the tools named
Daily counts for each companyHiring ActivityDaily open, new and closed postings for each company
Press releases and news itemsCompany NewsCompany announcements and regulatory disclosures classified into 13 event types
Scores built on the signalsIntent ScoresA weekly score from 0 to 100 for each company and topic, with its five strongest signals as evidence

Before you retire a spider, run both for a month on a sample of sites you know well. This query sets your detections beside the feed's, once your technology names are mapped to the feed's technology keys.

select
  coalesce(c.domain, f.domain)          as domain,
  coalesce(c.technology, f.technology)  as technology,
  c.first_detected_at                   as yours,
  f.first_seen_at                       as feed,
  f.seen_via
from my_crawler_detections c
full outer join company_technologies f
  on f.domain = c.domain
 and f.technology = c.technology
where coalesce(c.domain, f.domain) in (select domain from sample_sites);

Read the result in three groups. Rows on both sides confirm your mapping, and the two first-seen dates should sit close together. Rows found only on the feed's side are usually technologies detected in network requests or DNS records, which a page parser does not open. Rows found only on your side show what your spider reads that is particular to you, and they are the candidates for the one spider you keep. Take Acme Robotics, an illustrative company: the feed shows an email provider detected in its DNS records, with the date it was first seen, which your spider never looked for.

Fokals is delivered direct, by REST API and as bulk files in JSON, JSON Lines or CSV, which you load with your warehouse's own loader. The feeds of signals, website changes and postings run oldest first from a time you set, so the last cursor is your bookmark, as the guide to incremental API sync shows. Licensing is by written agreement for internal use, embedding in a product or redistribution. Other vendors license feeds of company data too, and the alternatives section of this library describes several.

Choosing by workload

  • License a feed for the rows every buyer needs in the same form: technologies on company websites, postings from company careers pages, announcements and the scores built on them. This is the route for a team that wants the data and not the crawl.
  • Keep a crawler for the pages and fields that are your own idea, and wherever the method is your product.
  • Add a managed service when your parsers work and the fetching does not.
  • Start from open data for reference facts and for research across very many pages.

The routes combine. A team can license the common rows, keep one spider for the pages that are its own idea and take reference facts from Wikidata. The comparison of building crawlers and licensing company data goes through the costs on each side.

Frequently asked questions

Should I build my own web crawler or license company data?

License when the rows are ones every buyer needs in the same form, such as technologies, job postings and announcements by company, and your engineers are better spent on what you do with them. Build when you need pages or fields that are particular to you, or when the collection method is itself what your product sells. A workload can split: a licensed feed for the common rows and one crawler for the rest.

What is the difference between Scrapy and Playwright?

Scrapy is an open-source Python framework for crawling websites and extracting structured data, with spiders, item pipelines and feed exports. Playwright is a browser automation library that drives Chromium, Firefox and WebKit through one API. They are used together: Scrapy's documentation recommends the scrapy-playwright integration for pages whose content its selectors cannot reach because scripts build it.

Does robots.txt give a crawler permission to read a site?

No. RFC 9309, which specifies the Robots Exclusion Protocol, sets out rules that crawlers are requested to honour and states that they are not a form of access authorisation. The file tells crawlers which paths a site owner asks them to stay out of, and it is not a grant of permission for the rest. Scrapy documents a setting, ROBOTSTXT_OBEY, that makes it respect those rules.

Can Common Crawl replace my own crawler?

For some work. Common Crawl publishes its crawls of the web as an open archive that is free for anyone to use, which suits research across a very large number of pages and saves you crawling them. Its data arrives as crawl archives, not as a daily reading of the pages you name, so watching specific companies day by day still needs your own crawler or a feed.

What does a licensed feed take off my engineering team?

The whole pipeline between a page and a finished record: fetching, the rules that decide what counts as a change, entity resolution, labelling, and the upkeep of each as websites change. Fokals delivers the result as dated records on one company ID, with daily and weekly datasets written once and never revised, labels under named, frozen versions and a documented schema. Your engineers load the data by REST API or bulk file and spend their time on what it is used for.

What stays with my team if I use a managed scraping service?

The rules that turn pages into dated rows for each company: what counts as a change, when a posting has closed, which company a website belongs to and how a row is labelled. Services such as Zyte API and Apify are documented as avoiding bans, automating a browser and running scrapers in a cloud. The decision about what to request, and from which sites, also stays with your team.

The queries and code on this page are examples to adapt. Test them in your own environment before you rely on them.

What this page says about the products it names was checked against their public documentation on 4 October 2026. Product and company names are trademarks of their owners. Fokals is not affiliated with them or endorsed by them.