Free tool
AI robots.txt builder
Pick the posture that matches your business, paste the file, and check it loads at your root.
Option 1 — Allow everything (recommended for most businesses)
If your website exists to bring you customers, this is almost certainly right.
User-agent: * Allow: / User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Bingbot Allow: / User-agent: Google-Extended Allow: / User-agent: Applebot-Extended Allow: / User-agent: CCBot Allow: / Sitemap: https://[yourdomain]/sitemap.xml
There is no Copilot crawler. Copilot answers come from Bing's index via Bingbot, so allowing Bingbot covers both. The real work for Copilot happens in Bing Webmaster Tools(opens in a new tab) — submit your sitemap there. See which crawlers to let in.
Option 2 — Retrieval yes, training no
A defensible middle path: appear in answers, opt out of training corpora. Note the boundary between the two blurs and the user-agent list changes, so revisit occasionally.
User-agent: * Allow: / # Live retrieval — we want to be cited User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / # Training crawlers — no User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / Sitemap: https://[yourdomain]/sitemap.xml
Option 3 — Block AI entirely
For paid archives, licensed content, or original work that is itself the product.
User-agent: * Allow: / User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: meta-externalagent Disallow: / Sitemap: https://[yourdomain]/sitemap.xml
Two things to know before you choose
Google-Extended does not affect Google Search. Blocking it removes you from Gemini-related uses only. Normal ranking is governed by Googlebot.
robots.txt is a request, not a lock. Well-behaved crawlers honour it. It offers no technical enforcement against ones that do not.
After you deploy
Open yoursite.com/robots.txt. Confirm plain text, no HTML, and that no earlier wildcard Disallow: / is sitting above your rules. Order matters less than most people think — the most specific matching user-agent group wins — but a stray blanket disallow catches everything that has no specific group.
Full background: which crawlers to let in.
Take this to your assistant
Paste it into ChatGPT, Copilot, Claude or Gemini and apply it to your own website.
Nothing is sent anywhere. The text is copied to your clipboard.