ddosNull AI Crawler & Scraper Blocker

Description

AI companies crawl the web to train their models, and some of them crawl hard. They take your content without asking, use up your server’s resources, and can slow your store down for real customers.

ddosNull AI Crawler & Scraper Blocker lets you decide, bot by bot, which AI crawlers get into your site. Everything is set up on activation. You don’t need to edit any files.

Features

  • 180+ known AI bots, each shown with its operator, purpose, crawl frequency, and whether it respects robots.txt.
  • One switch per bot to block or allow it. You can also block or allow every bot that matches a search or filter in one click.
  • Two layers of blocking:
    • Disallow rules are added to your robots.txt, so well-behaved crawlers stop visiting.
    • An HTTP 403 response goes to any request whose user agent matches a blocked bot. This also stops bots that ignore robots.txt.
  • Safe defaults. Major AI search engines, user-triggered assistants, and link-preview fetchers stay allowed, so your site still appears in ChatGPT, Perplexity, and Claude answers and in social link previews. Everything else is blocked. One click restores the recommended settings.
  • Bot list from the community-maintained ai.robots.txt project, bundled with the plugin and refreshed with each release. Optional daily updates (off by default) fetch new bots between releases. You choose what happens to newly discovered bots: block them unless they’re trusted, block them all, or allow them for manual review.
  • Blocked-request counters show which bots are hitting your site and when they last tried.
  • Search, filter, and sort by name, operator, category, robots.txt compliance, or number of blocked requests.
  • Fast. The check on each request is a single pattern match against a cached value. It adds no database queries to normal page views.
  • Light and dark mode.

Allowed by default

Blocking these services would hide your site from AI answers, shopping agents, or social shares, so they stay allowed unless you block them:

  • OpenAI: OAI-SearchBot, ChatGPT-User, ChatGPT Agent
  • Perplexity: PerplexityBot, Perplexity-User
  • Anthropic: Claude-SearchBot, Claude-User
  • Apple: Applebot (Siri, Spotlight, Safari). The AI-training opt-out, Applebot-Extended, is blocked.
  • Google: Google-Agent, GoogleAgent-URLContext. The AI-training opt-out, Google-Extended, is blocked.
  • DuckDuckGo: DuckAssistBot
  • Mistral: MistralAI-User
  • Amazon: Amzn-User, AmazonBuyForMe
  • Meta: facebookexternalhit (link previews on Facebook, WhatsApp, and Instagram), meta-externalfetcher

Search engines such as Googlebot and Bingbot are not AI crawlers and are never affected.

What user-agent blocking can’t do

This plugin stops bots that identify themselves. Many scrapers pretend to be a regular Chrome or Safari browser and rotate through thousands of IP addresses. The only way to catch those is by how they behave.

ddosNull Shield (free) detects bots by behavior and adds Layer-7 DDoS protection, without DNS changes. When both plugins are active, they work together:

  • Shield enforces your block and allow choices in its own firewall. In Shield’s Auto-Prepend mode, this happens before WordPress and your page cache load.
  • Your choices take priority over Shield’s built-in user-agent rules.
  • Allowed bots still pass through Shield’s IP reputation and DDoS checks. A scraper that pretends to be an allowed bot gains nothing.
  • Requests blocked by Shield still count toward the counters in this plugin.

External Services

This plugin can download the list of known AI crawlers from the ai.robots.txt project, which is hosted on GitHub. It only does so after an administrator turns on Automatic bot list updates in the plugin settings (off by default). Once enabled, the download runs once a day through WP-Cron, and also when you click the refresh icon on the settings page. With the setting off, the plugin makes no external requests.

  • URL requested: https://raw.githubusercontent.com/ai-robots-txt/ai.robots.txt/main/robots.json
  • Data sent: a plain HTTP GET request. No site data or personal data is included beyond what every HTTP request carries: your server’s IP address and WordPress’s user agent.
  • Provider: GitHub, Inc. See its Terms of Service and Privacy Statement.
  • The list is published under the MIT license by the ai.robots.txt contributors.

The plugin makes no other external requests.

Privacy

This plugin does not store visitors’ IP addresses or any other personal data. For each bot it keeps only a count of blocked requests and the time of the most recent one.

Source Code

The settings screen is a React app. Its uncompiled source code is included in the admin-src/ folder. To build it, run cd admin-src && npm install && npm run build.

Screenshots

Installation

  1. In your WordPress admin, go to Plugins Add New, search for “ddosNull AI Crawler & Scraper Blocker”, then click Install Now and Activate.
  2. Go to Settings AI Crawler Blocker.
  3. Protection is already on with the recommended settings. Review the bot list and adjust it if you like.

FAQ

Will this hurt my Google or Bing rankings?

No. Googlebot and Bingbot are not in the list. Google-Extended is only a robots.txt token that tells Google not to use your content for Gemini training. Blocking it does not affect Google Search.

Will my site still show up in ChatGPT, Perplexity, or Claude answers?

Yes, with the default settings. Their search crawlers and user-triggered fetchers are allowed. Only their model-training crawlers, such as GPTBot and ClaudeBot, are blocked.

What if a scraper fakes its user agent?

User-agent blocking only stops bots that identify themselves honestly. A scraper that pretends to be a browser needs behavioral detection, such as ddosNull Shield. Pretending to be one of the allowed AI bots gains a scraper nothing, because a browser user agent would already get it through.

Can blocked bots still read my robots.txt?

Yes, on purpose. robots.txt is always served, so compliant crawlers can read your rules and stop visiting.

I have a physical robots.txt file.

Your web server serves a physical robots.txt file directly, so WordPress can’t add rules to it. The settings page warns you when one exists. You can delete or rename the file, or copy the rules into it by hand. User-agent blocking works either way.

I use a page caching plugin or a CDN.

Page caches can serve a stored copy of a page before WordPress plugins run. In that case a blocked bot may receive the cached page, although the robots.txt rules still apply. With ddosNull Shield in Auto-Prepend mode, bots are blocked before the cache loads. Pages cached at a CDN edge (for example, Cloudflare’s “Cache Everything”) never reach your server, so no WordPress plugin can block those requests.

How often is the bot list updated?

A copy of the list ships with the plugin and is refreshed with every plugin release, so protection works from the first second without any external requests.

To get new bots between releases, turn on Automatic bot list updates under Settings AI Crawler Blocker. The plugin then downloads the list once a day through WP-Cron, and you can click the refresh icon to check right away. This setting is off by default.

What happens to a bot that is removed from the upstream list?

It stays in your list with your setting, labeled “Removed upstream”.

Can I change which bots are trusted?

Yes. Developers can adjust the list with the aicsb_trusted_bots filter. You can also block or allow any individual bot from the settings page.

What happens when I uninstall the plugin?

Uninstalling removes the plugin’s database table and all its settings. Deactivating keeps your settings and stops all blocking.

Reviews

There are no reviews for this plugin.

Contributors & Developers

“ddosNull AI Crawler & Scraper Blocker” is open source software. The following people have contributed to this plugin.

Contributors

Changelog

1.0.1

  • Renamed to ddosNull AI Crawler & Scraper Blocker.
  • Automatic bot list updates from GitHub are now opt-in and off by default. Without them, the bundled list is applied on each plugin update.
  • Sanitize request headers before matching.

1.0.0

  • Initial release.