Build an AI crawler policy for robots.txt
Create and inspect robots.txt rules for documented AI crawlers.
Choose crawler rules
Example: Sam at Acme Inc allows search and page fetches, blocks training crawlers, and keeps /private/ out of crawls.
“Block AI-specific crawlers” keeps broad Google, Bing, and Apple search crawling enabled. User-requested fetches may not honor these rules.
Documented crawler user-agents
Paths and sitemap
One path per line. Each path must start with /. Wildcards * and end marker $ are accepted.
robots.txt draft
Example is ready. Rules apply to compliant crawlers.
Example parsed: crawler-by-crawler access appears below.
A disallow rule does not keep a URL secret, secure private content, or remove an indexed result. Use access controls for private material and a noindex rule when you need to prevent indexing. Check the robots.txt your host or CDN serves; managed rules can change this draft.
The User-agent list is limited to names with operator or primary documentation. The control inspector is an estimate for common robots.txt grouping and path precedence, not a guarantee that every crawler will interpret rules the same way.
Embed this tool
Copy this snippet to show the tool on your site.
How it works
Each listed user-agent token and its described purpose links to documentation from its operator or a primary source, checked 28 September 2026.
The pasted-file inspector groups User-agent sections, checks named rules and wildcard rules, then applies the longest matching Allow or Disallow path rule. It is a practical preview, not a complete implementation of every crawler’s parser.
Cloudflare Content Signals are generated as optional policy lines. Cloudflare documents search, ai-input, and ai-train; its managed robots.txt feature can also add or change directives.
Sources and checked dates
- OpenAI crawler documentation
- Anthropic crawler documentation
- Perplexity crawler documentation
- Google crawler documentation
- Microsoft Bing crawler documentation
- Applebot documentation
- Meta crawler documentation
- DuckAssistBot documentation
- Common Crawl CCBot FAQ
- Cloudflare Content Signals
- Google robots.txt overview
Limits
- robots.txt is a voluntary crawl request. It does not block access, remove indexed URLs, or stop crawlers that ignore it.
- Some user-triggered fetchers may ignore robots.txt. WAF rules, authentication, and server controls are separate.
- Cloudflare may manage or append robots.txt rules for a zone; check the file your domain actually serves.
Common questions
- Does robots.txt block AI crawlers?
- It asks compliant crawlers not to fetch matching paths. It does not secure a page, guarantee crawler behavior, or remove a URL already in an index.
- What does Google-Extended control?
- Google documents Google-Extended as a control token for specified Gemini training and grounding uses. Googlebot still controls Google Search crawling.
- Do user-triggered fetchers follow these rules?
- Behavior varies. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it.
- What do Cloudflare Content Signals mean?
- They express site preferences for search indexing, real-time AI input, and model training. They do not guarantee that every crawler will honor them.
Related tools
- AI search readiness checker
Inspect a public page, its crawler rules and site files, then review a short list of fixes.
OperatorNest can take on repeat work, check with you before consequential steps, and leave a receipt for the result. See how an always-on operator works.
Hand off your first task tonight.
Tell us your email and what you'd hand off first. We'll send your access details and help you set up your operator.