Glossary

Web crawler

A web crawler visits pages automatically and records what it reads. How crawlers work, how they differ from scrapers and what separates responsible conduct from harmful.

Updated 5 October 20262 min read

A web crawler is a program that visits web pages automatically, follows the links it finds and records what it reads so that the content can be indexed or analysed. Search engines run crawlers to build their indexes, and so do data and research organisations that collect public information at scale.

How a crawler works

A crawler starts from seed addresses, fetches a page, extracts the links and the content it needs, adds new addresses to a queue and repeats, keeping track of what it has already seen. Crawling discovers pages; scraping pulls named fields from a page; most data collection does both. The quality of a crawler shows less in its speed than in its conduct: how it identifies itself, what it asks permission for and how little load it puts on each site.

What responsible conduct looks like

Five practices mark a responsible crawler:

  • It names itself in the request and links to a page that says what it reads and how to opt out.
  • It reads and obeys robots.txt, including crawl delays, before every request.
  • It makes one request at a time to a host, with a pause between requests.
  • It reads public pages only: no logins, no forms, no paywalls.
  • It does not change its identity to get past a block.

None of this is a legal opinion. The file robots.txt is a convention, not an access control, and whether a crawl is lawful depends on the jurisdiction, the site's terms and what is collected. It is the conduct a due diligence review usually asks about.

In Fokals data

Fokals delivers company data gathered from first-party company sources and public records, processed in-house, with every observation dated. The sourcing statement sets out the collection rules.

The rules a crawler obeys are set in robots.txt. Pages a company publishes about itself are a first-party data source, the record of where each value came from is data provenance, and a due diligence questionnaire asks how a vendor crawls.

Frequently asked questions

What is the difference between a web crawler and a web scraper?

A crawler discovers pages by following links and keeps a record of what it has visited. A scraper extracts specific fields, such as a price or a job title, from a page it already knows. The terms overlap in practice because most data collection does both: it crawls to find pages, then scrapes the fields it needs from them. Responsible conduct applies to both.

Is it legal to crawl a website?

There is no single answer. It depends on the country, the site's terms of use, how the crawl is done and what is collected, above all whether personal data is involved. Identifying the crawler, obeying robots.txt, reading public pages only and collecting no personal data are the points a review usually starts from, but they describe conduct and are not legal advice. Take advice for your own case.

How do websites control crawlers?

Through several layers. A robots.txt file asks crawlers which paths to avoid and how fast to go. Rate limits and network protection slow or block requests that look automated. Logins and paywalls keep content from anyone without an account. Terms of use set the rules for reuse. Robots.txt works only for crawlers that choose to obey it, while rate limits, network protection and logins act on every request whatever the crawler intends.