MonetizationOS Blog

What to Do When AI Crawlers Ignore Your robots.txt

Guides
July 21, 2026
3 minutes min read
What to Do When AI Crawlers Ignore Your robots.txt
In this article
  • 1
    Introduction

Robots.txt depends on crawlers choosing to obey it. The problem is most of them are choosing not to. It’s a polite request, honored by convention, and the convention is breaking down.

The traffic itself has changed shape: more than half of web traffic is now automated — 53% in 2025, up from 51% in 2024, in the traffic Imperva measured for its 2026 Bad Bot Report.

So the honest starting point is this.

If your machine-traffic strategy is a robots.txt file, you do not have a strategy.

Blocking works, but it also costs you.

The obvious response is to block harder: firewall rules, IP blocklists, server-level rejections. These do something robots.txt cannot, because they refuse the request instead of asking nicely.

But blocking is not free.

A December 2025 study by Hangcheng Zhao (Rutgers) and Ron Berman (Wharton) found that large news publishers blocking generative-AI bots saw roughly 23% lower total traffic — and 14% lower human traffic — than comparable publishers that did not block. Block everything and you do not simply keep the machines out. For large publishers, blanket blocking may also reduce the human traffic arriving through AI-mediated discovery.

It is also a fight that gets harder every quarter.

Cloudflare removed Perplexity from its verified-bot list after finding crawlers that rotated IP addresses and disguised themselves as Chrome on macOS to get around firewall rules. Perplexity disputed Cloudflare’s account.

When a crawler can look like a browser, detection alone stops being a strategy.

Govern instead: sort the content, set the rule

Between “block everything” and “allow everything” sits the thing that actually protects the asset — which is a policy.

Not every piece of content deserves the same treatment. A workable policy starts by separating content into a few classes.

Each class gets one clear rule:

  • Archive and evergreen. Allow access under terms, or hold for a window and then open it.
  • Breaking news. Protect it while human traffic is most valuable, then release.
  • Subscriber-only analysis and investigations. The highest-value tier. Restrict hard, or open only to partners on agreed terms.
  • Structured data and proprietary datasets. Treat these as distinct commercial assets, with explicit terms, identification and usage tracking.

The exact taxonomy matters less than the principle.

“Block everything” and “allow everything” are both abdications.

Between them sits a set of commercial decisions, and decisions need infrastructure that can carry them out on every request, at the speed machine traffic moves.

This is an access decision, not a security setting

Detection still matters, because you cannot govern what you cannot identify. But detection answers “is this a bot?” The question that decides what happens to your content is “should this bot get this, and on what terms?”

That is a policy and enforcement decision, informed by identity and commercial terms, and it belongs at the edge, before content is served.

It’s the work MonetizationOS exists to do: govern access for every visitor, human or machine, through one layer that decides at the point of request.

AI systems are already intermediaries between publishers and some readers, while their crawlers form part of a machine-traffic landscape that now accounts for more than half of measured web traffic.

The practical question is now less about blocking, and more about which machines should reach which content, under what terms, and how those terms are enforced.

No items found.

Get started with instant momentum

Take full control of your intellectual property with a fast, future-ready monetization engine.

Get Started for free