ROBOTS / SEARCH & TRAINING
Bulk AI Crawler Policy Tester
Test up to ten paths against robots.txt for search, training and user-requested retrieval agents. See the matched rule instead of a guessed allow/block label.
A PUBLIC PAGE / A USEFUL ANSWER
Check your crawler policy.
Your findings will appear here.
Run a check to see measured values, matched rules and specific next steps.
THE PROCESS
A few steps to a useful result.
- 01
Choose one public HTTPS site
Each hostname has its own robots.txt. Enter one site and up to ten absolute paths, one per line. Query strings and case can affect matching.
- 02
Read the matched policy
The checker evaluates agent groups, wildcard and end-anchor rules, and longest-path matches. Equal-length Allow rules win over Disallow rules.
- 03
Keep crawler roles separate
Search, training and user retrieval can have independent controls. Robots permission does not prove a verified crawler can pass your firewall, that a page is indexed or that a change is already applied.
GOOD QUESTIONS
Before you use the results.
Can I block training and allow search?
Some operators document separate search and training agents, including OpenAI and Anthropic. Configure each relevant bot and hostname, then check the operator's current documentation.
Does an allowed path guarantee indexing or citations?
No. Robots permission is one piece of access policy. Indexing, citations, genuine bot identity and firewall permission require separate evidence.
OFFICIAL GUIDANCE / REVIEWED 2026-10-02
Know where the rules come from.
Report results reflect the response available when you run the check. Operator requirements and private account settings can change.