Skip to content
OperatorNest

Build an AI crawler policy for robots.txt

Create and inspect robots.txt rules for documented AI crawlers.

Example: Sam at Acme Inc: Search and user-requested fetchers are allowed, GPTBot and ClaudeBot are blocked, and /private/ is excluded.

Choose crawler rules

Example: Sam at Acme Inc allows search and page fetches, blocks training crawlers, and keeps /private/ out of crawls.

Google robots.txt specification, checked 28 September 2026

“Block AI-specific crawlers” keeps broad Google, Bing, and Apple search crawling enabled. User-requested fetches may not honor these rules.

Documented crawler user-agents

Paths and sitemap

One path per line. Each path must start with /. Wildcards * and end marker $ are accepted.

Optional Cloudflare Content Signals

These express policy preferences. They are not crawler access rules.

Cloudflare Content Signals documentation, checked 28 September 2026

robots.txt draft

Example is ready. Rules apply to compliant crawlers.

Example parsed: crawler-by-crawler access appears below.

    A disallow rule does not keep a URL secret, secure private content, or remove an indexed result. Use access controls for private material and a noindex rule when you need to prevent indexing. Check the robots.txt your host or CDN serves; managed rules can change this draft.

    The User-agent list is limited to names with operator or primary documentation. The control inspector is an estimate for common robots.txt grouping and path precedence, not a guarantee that every crawler will interpret rules the same way.

    Embed this tool

    Copy this snippet to show the tool on your site.

    How it works

    Each listed user-agent token and its described purpose links to documentation from its operator or a primary source, checked 28 September 2026.

    The pasted-file inspector groups User-agent sections, checks named rules and wildcard rules, then applies the longest matching Allow or Disallow path rule. It is a practical preview, not a complete implementation of every crawler’s parser.

    Cloudflare Content Signals are generated as optional policy lines. Cloudflare documents search, ai-input, and ai-train; its managed robots.txt feature can also add or change directives.

    Sources and checked dates

    Limits

    • robots.txt is a voluntary crawl request. It does not block access, remove indexed URLs, or stop crawlers that ignore it.
    • Some user-triggered fetchers may ignore robots.txt. WAF rules, authentication, and server controls are separate.
    • Cloudflare may manage or append robots.txt rules for a zone; check the file your domain actually serves.

    Common questions

    Does robots.txt block AI crawlers?
    It asks compliant crawlers not to fetch matching paths. It does not secure a page, guarantee crawler behavior, or remove a URL already in an index.
    What does Google-Extended control?
    Google documents Google-Extended as a control token for specified Gemini training and grounding uses. Googlebot still controls Google Search crawling.
    Do user-triggered fetchers follow these rules?
    Behavior varies. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it.
    What do Cloudflare Content Signals mean?
    They express site preferences for search indexing, real-time AI input, and model training. They do not guarantee that every crawler will honor them.

    OperatorNest can take on repeat work, check with you before consequential steps, and leave a receipt for the result. See how an always-on operator works.

    Hand off your first task tonight.

    Tell us your email and what you'd hand off first. We'll send your access details and help you set up your operator.