AI Crawlers Are Hitting Your Site Right Now. Should You Block Them?
DEV Community

AI Crawlers Are Hitting Your Site Right Now. Should You Block Them?

Go check your server logs. I'll wait. See those user agents you don't recognize? GPTBot, ClaudeBot, CCBot, PerplexityBot, Bytespider, Amazonbot - that's not Googlebot doing its usual rounds. That's a second, newer wave of crawlers: bots collecting your content to train AI models and feed retrieval systems that answer questions without ever sending you a visitor. As developers, we've spent a decade optimizing for search crawlers. Now we have to make a call our SEO colleagues can't make for us: do we let these bots in, or shut the door? Here's how to think about it like an engineer.

What AI crawlers actually do

Search crawlers (Googlebot, Bingbot) fetch your pages to build an index that sends you traffic. The deal is obvious: they crawl, you get visitors. AI crawlers have a murkier value exchange. They generally do one of two things:

  • Training data collection - scraping the web to train the next generation of language models. Your content becomes weights in a model. You get nothing directly.
  • Retrieval for AI answers - fetching pages to ground responses in tools like AI Overviews, ChatGPT, or Perplexity. Here there is a potential upside: being cited in an AI answer is the new front page.

The tricky part? The same bot often does both, and the operators are not always transparent about which. That's why a blanket "block everything" or "allow everything" is the wrong move.

The costs developers actually feel

This isn't abstract. AI crawlers are aggressive - some ignore crawl-delay, hit your site thousands of times a day, and request heavy pages (image galleries, faceted search URLs, API endpoints). I've seen them:

  • Spike hosting bills on metered infrastructure - one misbehaving bot can cost real money on serverless.
  • Pollute analytics - bot traffic mixed into your conversion data unless you filter it.
  • Trigger rate limits and WAF rules - ironically, your own defenses can start blocking legitimate users when a bot hammers shared IPs.
  • Scrape behind-login or paywalled content via aggressive URL guessing on sites with predictable routes.

If your site is small, a single aggressive crawler can be a noticeable percentage of your total traffic. Check before you decide anything.

The case for letting them in

Blocking feels good until you realize what you're giving up. AI-powered answer engines are where a growing share of discovery happens. If your content is never crawled by these systems, it can never be cited by them. For documentation sites, blogs, and open knowledge bases, that's a genuine distribution channel you're closing. There's also a middle path most teams miss: you don't have to treat every AI crawler the same. Retrieval bots that cite sources (and can send traffic) are very different from pure training scrapers. A selective policy beats a binary one.

Make the decision like an engineer

Here's the framework I use:

  1. Measure first. Pull 7 days of logs. Which AI user agents hit you, how often, which URLs, how much bandwidth? No data, no decision.
  2. Classify your content. Public marketing content and docs? Probably fine to allow retrieval. User-generated content, paywalled material, or expensive-to-render pages? Lean toward blocking.
  3. Separate training from retrieval. Allow bots that attribute and cite; block pure scrapers that give nothing back.
  4. Start permissive, tighten deliberately. It's easier to block a specific bad actor you can name than to undo months of invisibility in AI answers.

For a deeper breakdown of the trade-offs - including which specific bots do what - this guide on whether you should block AI crawlers walks through each major crawler's behavior and the robots.txt directives that control them.

Controlling them properly (not just vibes)

If you decide to block or throttle, do it at the right layer:

  • robots.txt - the standard mechanism. But write it carefully: a sloppy Disallow can nuke your search traffic too.
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: ClaudeBot
Disallow: /

Keep Googlebot and other search crawlers on their own rules, and test the file - one misplaced wildcard has taken down more sites than any algorithm update.

  • Server-level controls - for bots that ignore robots.txt (and some do), enforce at the edge: rate-limit by user agent, return 403s, or challenge with your WAF. robots.txt is a request, not a lock.
  • llms.txt - the emerging convention: a markdown file at /llms.txt that gives AI systems a clean, structured summary of your site. Think of it as a sitemap for language models. Early, but worth watching.

Why this matters beyond SEO: the RAG connection

Here's the part most web developers miss: AI crawlers aren't just an SEO problem. The content they collect feeds enterprise RAG systems - retrieval-augmented generation pipelines that companies build on top of web-scale data. Your documentation, your tutorials, your forum answers: they end up chunked, embedded, and retrieved inside someone else's AI product. That cuts both ways. If you build AI products, you should understand exactly how this pipeline consumes web content - because the same crawling, chunking, and embedding mechanics apply to your own knowledge bases. The crawler problem and the RAG problem are the same problem viewed from opposite sides.

And then there are the agents

The next wave is already here: AI agents that browse the web like users - clicking, scrolling, filling forms. They don't respect robots.txt the way crawlers do because they often drive real browsers. If you thought bot management was solved, agents are about to reopen the whole discussion: how do you distinguish a helpful agent completing a task for a user from a scraper wearing a browser costume? For now: get your crawler policy right, instrument your traffic so you can see what's actually hitting you, and keep the policy reviewable. The bot landscape is changing quarterly.

The checklist

  • [ ] Audit 7 days of logs for AI user agents and their bandwidth cost
  • [ ] Classify content: public / attributable vs. private / expensive
  • [ ] Write explicit robots.txt rules per bot family - don't blanket-block
  • [ ] Enforce at the edge for bots that ignore robots.txt
  • [ ] Revisit quarterly - new bots appear constantly

The bots aren't going away. The developers who measure, decide deliberately, and control at the right layer will be fine. The ones running on vibes and blanket blocks will either pay the bandwidth bill or vanish from AI answers entirely. Pick your trade-off on purpose.

More on crawlers, robots.txt, and technical SEO: browse the SEO and AI tags here on DEV.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.