Skip to content
Directory/Methodology
How the registry works

How we track, verify, and rate every bot

Every entry in the registry follows the same process - from the moment a crawler first shows up in a data source to the sentiment rating you see on its page. Here is exactly how that works, and how to tell us when we get something wrong.

Overview

Bot WHOIS is an independent registry of automated web clients - AI crawlers, search engines, SEO tools, scrapers, monitors, and more. We aggregate signals from several independent sources, cross-reference them against operator-published records, and enrich each entry before it is listed. No operator pays to be included or to change its rating.

642
Bots tracked
Multi-source
Cross-referenced
Automated
Re-sync
Open
Corrections
Step 1

How we collect bots

A bot enters the registry from one of four kinds of source. Every candidate is de-duplicated by name and user-agent before it is enriched and listed.

A

Cloudflare Radar

We ingest Cloudflare Radar's verified-bot list and traffic data, which supplies the authoritative identity, category, and relative request-volume signal for a large share of the registry.

B

Crawler databases & operator docs

We pull from public crawler databases and the official user-agent and IP-range files operators publish, so a bot can be listed soon after its operator announces it.

C

Curated allow-lists

We cross-reference community-maintained crawler allow-lists (such as the widely-used GitHub lists) to catch well-known bots that our other sources miss.

D

Live discovery

Unknown user-agents seen in real request logs surface as candidates, which are researched and only published once they clear a confidence threshold.

Step 2

Enrichment

Once a bot is identified, we enrich its profile from additional sources: a plain-language summary of what it does and why it crawls, its published IP ranges, block and allow guidance, recent community discussion, relevant news, and a search-interest trend. Enrichment runs continuously so profiles deepen over time.

Step 3

Verification

A bot is marked Verifiedwhen an authoritative source - such as Cloudflare's verified-bot programme or an operator's own published records - confirms it is genuinely who it claims to be. Everything else is labelled Community-sourced: still listed, but not yet independently confirmed.

Why it matters.A user-agent string is trivial to spoof. Tying an entry back to an operator's own published identity and IP ranges is what separates a real operator bot from an impersonator.
Step 4

Sentiment rating

Each bot gets one of three ratings. A rating combines how well we can verify the bot's identity with any behaviour our sources have flagged - it is not a guarantee that we have tested the bot on your site. Always confirm against your own logs before allowing or blocking.

GoodBot

A verified, well-identified crawler from a known operator, with nothing flagged against it by our sources.

Undecided

Identified but not independently verified, or with open questions - review the details before deciding.

BadBot

Flagged by our sources for abusive behaviour - ignoring robots.txt, disguising its identity, or crawling aggressively enough to degrade a site.

Step 5

Categorisation

Every bot is filed under one primary category based on what its traffic is for - training data, live assistance, search indexing, SEO tooling, scraping, monitoring, webhooks, and so on. Categories are kept coarse so the directory stays navigable; the full taxonomy with a live bot count for each lives on the Categories page.

Ongoing

Updates & corrections

The registry re-syncs on an automated schedule: sources are re-fetched, new bots are picked up, and existing profiles are re-enriched. Every bot page shows its last-updated date.

If you believe a rating or fact is wrong - especially if you operate the bot - tell us and we will review it against our sources and update the entry.