Two pipelines feed one list. The blocklist is derived from our main pre-categorized database of 102 million domains — kept current by checking and classifying roughly 300,000 newly registered domains every day — plus a dedicated discovery pipeline that monitors the sources where AI tools appear first. Both run daily.
Everything begins with our main pre-categorized database of 102 million domains — the same commercial database that powers websitecategorizationapi.com. Every domain in the corpus is classified by function, content type, and risk category.
102 million domains already classified by function, content type, and risk category form the starting point. The blocklist is never assembled from scratch or from someone's bookmarks.
To keep the database current, roughly 300,000 newly registered domains from zone files, CT logs, and registration feeds are checked and classified every day. Many turn out to be empty or parked — those don't inflate the database.
Any domain in the corpus — existing or newly classified — that is identified as an AI tool automatically enters the AI-tools feed. This is how we catch AI tools that exist as domains but appear in no directory, no press coverage, and no curated list.
Classified AI-tool domains flow into the shared delivery step below, where they merge with the output of the discovery pipeline.
Separately from the 102M database, we run a dedicated pipeline that monitors sources specifically frequented by AI tools.
App directories, product-launch platforms, developer communities, academic repositories, and the open web — scanned daily, with overlapping sources by design.
Candidate domains are classified into the 18 categories, with tool name and subcategory assignment.
Dead domains are pruned. Aliases and subdomains are resolved. What remains is a clean, deduplicated set.
New AI-tool domains are merged into the shared delivery step within the same daily cycle. No manual list can match this cadence.
Both pipelines run every day, against a corpus and a source set far beyond what any curated bookmark list covers.
Both pipelines merge into a single feed. The result is exported every day to CSV, JSON, EDL, PAC, hosts, and DNS RPZ formats, and served through the REST API — so your enforcement points are never more than 24 hours behind reality.
Root domain (e.g. claude.ai)
Human-readable product name
One of 18 categories and 172 subcategories
Live, redirected, or parked
210-language detection
Each domain lands in exactly one of 18 functional categories — from Agents & Automation to Video — and one of 172 subcategories, so your policies stay unambiguous: allow approved code assistants, block consumer chatbots, monitor the rest. Full definitions with example domains live on the taxonomy page.
We walk enterprise evaluators through both pipelines — the 102M-domain corpus extraction and the daily discovery cycle. Tell us your evaluation criteria.
Tell us about your evaluation criteria and we will schedule a technical walkthrough of our classification methodology.