AI Tools Blocklist
Home AI Tools Database Taxonomy Pricing
Solutions
Enterprise IT & CISO Education Firewall Admins Shadow AI Prevention REST API Customer Login
Download Free Sample
Classification Methodology

How We Classify 17,410+ AI Tools
from a 102M-Domain Corpus

Two pipelines feed one list. The blocklist is derived from our main pre-categorized database of 102 million domains — kept current by checking and classifying roughly 300,000 newly registered domains every day — plus a dedicated discovery pipeline that monitors the sources where AI tools appear first. Both run daily.

102MDomain Corpus
17,410+AI Tools Classified
18Functional Categories
102M-Domain Corpus
~300K Domains Checked Daily
Two Daily Pipelines
18 Categories
172 Subcategories
Daily Exports
Pipeline A

Pipeline A — From the 102M-Domain Database

Everything begins with our main pre-categorized database of 102 million domains — the same commercial database that powers websitecategorizationapi.com. Every domain in the corpus is classified by function, content type, and risk category.

102M Corpus
Daily Refresh
AI-Tool Extraction
Delivery
A1

102M-Domain Corpus

102 million domains already classified by function, content type, and risk category form the starting point. The blocklist is never assembled from scratch or from someone's bookmarks.

A2

Daily Refresh

To keep the database current, roughly 300,000 newly registered domains from zone files, CT logs, and registration feeds are checked and classified every day. Many turn out to be empty or parked — those don't inflate the database.

A3

AI-Tool Extraction

Any domain in the corpus — existing or newly classified — that is identified as an AI tool automatically enters the AI-tools feed. This is how we catch AI tools that exist as domains but appear in no directory, no press coverage, and no curated list.

A4

To Delivery

Classified AI-tool domains flow into the shared delivery step below, where they merge with the output of the discovery pipeline.

Pipeline B

Pipeline B — AI-Tool Discovery

Separately from the 102M database, we run a dedicated pipeline that monitors sources specifically frequented by AI tools.

B1

Source Monitoring

App directories, product-launch platforms, developer communities, academic repositories, and the open web — scanned daily, with overlapping sources by design.

B2

Classification

Candidate domains are classified into the 18 categories, with tool name and subcategory assignment.

B3

Deduplication

Dead domains are pruned. Aliases and subdomains are resolved. What remains is a clean, deduplicated set.

B4

To Delivery

New AI-tool domains are merged into the shared delivery step within the same daily cycle. No manual list can match this cadence.

Why a second pipeline? It finds AI tools that may not yet have enough web presence for the corpus classifiers to catch — tools announced yesterday, in beta, or serving a niche vertical. They're identified, classified, and added to the feed within the same daily cycle.
The data underneath

Scale that no manual list can match

Both pipelines run every day, against a corpus and a source set far beyond what any curated bookmark list covers.

102MDomains in the corpus
~300KNew domains checked daily
17,410+AI-tool domains classified
24hMaximum feed age
Zone files CT logs Registration feeds App directories Product-launch platforms Developer communities Academic repositories Open web
Delivery

One Unified, Deduplicated Feed — Exported Daily

Both pipelines merge into a single feed. The result is exported every day to CSV, JSON, EDL, PAC, hosts, and DNS RPZ formats, and served through the REST API — so your enforcement points are never more than 24 hours behind reality.

Domain

Root domain (e.g. claude.ai)

Tool Name

Human-readable product name

Category

One of 18 categories and 172 subcategories

Status

Live, redirected, or parked

Language

210-language detection

Taxonomy

18 Categories, 172 Subcategories

Each domain lands in exactly one of 18 functional categories — from Agents & Automation to Video — and one of 172 subcategories, so your policies stay unambiguous: allow approved code assistants, block consumer chatbots, monitor the rest. Full definitions with example domains live on the taxonomy page.

Corrections
Dual-purpose domains are assigned to their primary function so you can make an informed call. If you find a misclassification, report it to [email protected] and it is corrected in a subsequent daily export.

Request a Technical Deep Dive

We walk enterprise evaluators through both pipelines — the 102M-domain corpus extraction and the daily discovery cycle. Tell us your evaluation criteria.

Request a Technical Deep Dive

Tell us about your evaluation criteria and we will schedule a technical walkthrough of our classification methodology.