Free AI visibility check. One page, all four dimensions. No signup, nothing installed. Check a page →
Home/Blog

Commentary

The honest state of AI crawler blocking

Large publishers have good reasons to block. Most small businesses are copying a decision that was never made for them.

6 min readVirtual Planet
robots.txt one line each GPTBot ClaudeBot Bingbot Your pages

A significant share of large news sites now block GPTBot, ClaudeBot and their equivalents in robots.txt. That is a rational decision for a business whose product is the text itself. It is a strange decision to copy if your product is plumbing.

Why publishers block

Their content is the inventory. If an assistant summarises the article, the reader never arrives, the ad never loads, and the subscription never converts. Blocking is a negotiating position as much as a technical one — several major outlets have signed licensing deals precisely because blocking gave them something to trade.

That logic is sound and specific. It depends on the content being the product.

Why most businesses are not in that position

If you sell a service, your website is not the inventory. It is the shop window. An assistant summarising your opening hours and recommending you is not taking a sale away — it is the sale.

The publisher's problem is substitution. Your problem is obscurity. Those call for opposite responses, and it is worth being clear which one you have before you copy someone else's robots.txt.

The genuinely difficult middle

Some businesses sit awkwardly between the two. If you have spent a decade building a library of technical guides that brings in leads, an assistant that answers from your guides without sending anyone your way is a real cost.

I do not think there is a clean answer there. What I would say is that the decision should be made deliberately, per crawler, with an understanding of what each one does — training, retrieval, or live search — rather than by pasting a block list found on a forum. Our crawler guide lists what each known agent is actually for.

The thing people get wrong mechanically

Blocking a crawler in robots.txt does not remove you from the model. Training runs have already happened. What it changes is whether an assistant can fetch your page now to answer a question about you accurately — which means the practical effect of blocking is often that assistants describe you from stale or second-hand information rather than not at all.

That is close to the worst outcome available: still discussed, no longer accurate, and unable to correct it.

What I would do

For most service businesses: allow everything, make the site readable, and treat assistants as a distribution channel you did not have to pay for.

For publishers and anyone whose text is the product: block deliberately, know what you are trading, and revisit it when the licensing market matures.

For everyone: check what your robots.txt currently says, because a surprising number of sites are blocking crawlers they never intended to block, inherited from a theme or a plugin default.

Take this to your assistant

Paste it into ChatGPT, Copilot, Claude or Gemini and apply it to your own website.

Nothing is sent anywhere. The text is copied to your clipboard.