{"id":366338,"date":"2026-09-19T16:48:18","date_gmt":"2026-09-19T16:48:18","guid":{"rendered":"https:\/\/wordpress.org\/plugins\/citely-ai-crawler-controls\/"},"modified":"2026-09-19T21:11:38","modified_gmt":"2026-09-19T21:11:38","slug":"crawlune-ai-crawler-controls","status":"publish","type":"plugin","link":"https:\/\/wordpress.org\/plugins\/crawlune-ai-crawler-controls\/","author":23539523,"comment_status":"closed","ping_status":"closed","template":"","meta":{"version":"0.62.6","stable_tag":"0.62.6","tested":"7.1.1","requires":"6.4","requires_php":"8.1","requires_plugins":null,"header_name":"Crawlune \u2013 AI Crawler Controls","header_author":"WP Shelf","header_description":"AI visibility for WordPress: llms.txt, AI crawler controls, and an audit of whether your content can be quoted by answer engines.","assets_banners_color":"594cb0","last_updated":"2026-09-19 21:11:38","external_support_url":"","external_repository_url":"","donate_link":"","header_plugin_uri":"","header_author_uri":"","rating":0,"author_block_rating":0,"active_installs":0,"downloads":38,"num_ratings":0,"support_threads":0,"support_threads_resolved":0,"author_block_count":0,"sections":["description","changelog"],"tags":{"0.62.5":{"tag":"0.62.5","author":"wpshelf","date":"2026-09-19 16:48:08","revision":3703487},"0.62.6":{"tag":"0.62.6","author":"wpshelf","date":"2026-09-19 21:11:38","revision":3703673}},"upgrade_notice":[],"ratings":[],"assets_icons":{"icon-128x128.png":{"filename":"icon-128x128.png","revision":3703487,"resolution":"128x128","location":"assets","locale":"","width":128,"height":128},"icon-256x256.png":{"filename":"icon-256x256.png","revision":3703487,"resolution":"256x256","location":"assets","locale":"","width":256,"height":256},"icon.svg":{"filename":"icon.svg","revision":3703487,"resolution":false,"location":"assets","locale":false}},"assets_banners":{"banner-1544x500.png":{"filename":"banner-1544x500.png","revision":3703487,"resolution":"1544x500","location":"assets","locale":"","width":1544,"height":500},"banner-772x250.png":{"filename":"banner-772x250.png","revision":3703487,"resolution":"772x250","location":"assets","locale":"","width":772,"height":250}},"assets_blueprints":{},"all_blocks":[],"tagged_versions":["0.62.5","0.62.6"],"block_files":[],"assets_screenshots":{"screenshot-1.png":{"filename":"screenshot-1.png","revision":3703487,"resolution":"1","location":"assets","locale":"","width":1216,"height":950},"screenshot-2.png":{"filename":"screenshot-2.png","revision":3703487,"resolution":"2","location":"assets","locale":"","width":1216,"height":950},"screenshot-3.png":{"filename":"screenshot-3.png","revision":3703487,"resolution":"3","location":"assets","locale":"","width":1216,"height":950}},"screenshots":{"1":"AI visibility \u2014 the median extractability score, what the score is made of, and how many pages are missing each signal","2":"llms.txt \u2014 what is published, what is held back and why","3":"Settings \u2014 licensing, e-mailed reports and opt-in usage reporting"}},"plugin_section":[],"plugin_tags":[2353,247966,244604,1117,186],"plugin_category":[55],"plugin_contributors":[281598,274470],"plugin_business_model":[],"class_list":["post-366338","plugin","type-plugin","status-publish","hentry","plugin_tags-ai","plugin_tags-gptbot","plugin_tags-llms-txt","plugin_tags-schema","plugin_tags-seo","plugin_category-seo-and-marketing","plugin_contributors-matik","plugin_contributors-wpshelf","plugin_committers-wpshelf"],"banners":{"banner":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/banner-772x250.png?rev=3703487","banner_2x":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/banner-1544x500.png?rev=3703487","banner_rtl":false,"banner_2x_rtl":false},"icons":{"svg":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/icon.svg?rev=3703487","icon":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/icon.svg?rev=3703487","icon_2x":false,"generated":false},"screenshots":[{"src":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/screenshot-1.png?rev=3703487","caption":"AI visibility \u2014 the median extractability score, what the score is made of, and how many pages are missing each signal"},{"src":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/screenshot-2.png?rev=3703487","caption":"llms.txt \u2014 what is published, what is held back and why"},{"src":"https:\/\/ps.w.org\/crawlune-ai-crawler-controls\/assets\/screenshot-3.png?rev=3703487","caption":"Settings \u2014 licensing, e-mailed reports and opt-in usage reporting"}],"raw_content":"<!--section=description-->\n<p>Crawlune audits how easily your published content can be extracted, controls AI\ncrawler access, and publishes machine-readable summaries. These features run\nlocally and do not require an account or license. It does not measure or promise\ncitations by answer engines.<\/p>\n\n<p>Implemented:<\/p>\n\n<ul>\n<li>Per-agent AI crawler controls for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended, distinguishing training crawlers from retrieval crawlers<\/li>\n<li>Three named positions to start from \u2014 Open, \"Answers yes, training no\", Closed \u2014 each setting every crawler and the matching content signals in one go<\/li>\n<li>robots.txt directives generated from those rules<\/li>\n<li>Optional refusal (403) for a crawler you blocked that requests pages anyway \u2014 off by default, and the requests keep being counted<\/li>\n<li>A count of people arriving from ChatGPT, Claude, Perplexity, Gemini and Copilot \u2014 click-throughs, which is a floor under citations rather than a measure of them<\/li>\n<li>Content signals (search, ai-input, ai-train) in robots.txt, which state what a page may be used for rather than who may fetch it \u2014 off until you set them, and a declaration rather than an enforcement<\/li>\n<li>The same refusal of AI training repeated as a noai response header, a robots meta tag and a W3C TDM reservation at \/.well-known\/tdmrep.json<\/li>\n<li>Markdown for agents: every page also available at its own address with .md on the end, and on request via Accept: text\/markdown \u2014 off by default, and honouring the same exclusions as llms.txt<\/li>\n<li>IndexNow submission when you publish or update a page \u2014 off by default, and honouring the same exclusions as llms.txt<\/li>\n<li>An extractability audit scoring content on liftable answer blocks, question-shaped headings, external citations, quotable statistics, modification dates and thin content<\/li>\n<li>A site-wide visibility report aggregating that across all published content, worst pages first<\/li>\n<li>\/llms.txt and \/llms-full.txt, generated from your real content and refreshed when you publish<\/li>\n<li>Admin screens for the report, a per-page work-list behind every number, crawler controls, llms.txt and licensing, plus REST routes for the same<\/li>\n<li>Background analysis of your content, so none of the above is computed while you wait for a page to load<\/li>\n<li>Structured data for your pages when no SEO plugin is already providing it<\/li>\n<li>AI rewriting of a page's opening into a liftable answer block, reviewed as a word-level diff before anything is applied<\/li>\n<li>AI rewriting of section headings toward the question they answer, keeping each heading's level and its existing link address<\/li>\n<li>Generated FAQ data for pages that answer questions they never ask, where every proposed answer has to be text already on the page \u2014 anything else is discarded before you see it<\/li>\n<li>Applying a rewrite without disturbing surrounding block markup, and reverting a run from a journal<\/li>\n<li>License activation against the WP Shelf licensing service<\/li>\n<\/ul>\n\n<p>Hosted rewriting requires a WP Shelf license and an enabled generation service.\nHosted rewriting is enabled for activated licenses. Public licensing enrollment\nis not yet open. Local audits and crawler tools do not\ndepend on it.<\/p>\n\n<p>Not included:<\/p>\n\n<ul>\n<li>Citation tracking \u2014 whether a given engine actually quoted you. Counting people who click through from an assistant is implemented and is a different thing: it is a floor under citations, not a measure of them<\/li>\n<\/ul>\n\n<p>Crawlune does not claim it will get your site cited by ChatGPT, Perplexity, Claude\nor Google AI Overviews. That is not measurable from inside WordPress and will not\nbe claimed. What it reports is what it measured about your content.<\/p>\n\n<h3>External services<\/h3>\n\n<p>The content audit, crawler controls, structured-data checks and generated\nllms.txt files run on your site. External requests occur only for the optional\nfeatures described below. Activating the plugin alone does not opt you into\nusage reporting or start generation.<\/p>\n\n<p><strong>1. IndexNow.<\/strong> With \"Submit a page to IndexNow\" on, publishing or updating a\npage sends that page's address to <code>https:\/\/api.indexnow.org\/indexnow<\/code>, together\nwith your site's host name and a verification key this site publishes at its own\nroot. Nothing else goes with it \u2014 no content, no e-mail address, no license key.\nIndexNow is a shared endpoint operated by Microsoft on behalf of Bing, Yandex,\nSeznam and Naver; the same URL is submitted to all of them once. At most one\nsubmission per address per hour. Service information and terms:\nhttps:\/\/www.indexnow.org\/faq\nPrivacy information: https:\/\/www.microsoft.com\/privacy\/privacystatement<\/p>\n\n<p><strong>2. WP Shelf licensing and generation.<\/strong> A WP Shelf account and a\nproduct-specific license are required to use the hosted generation service;\nthe local features above remain free without one.<\/p>\n\n<ul>\n<li>Service host: <code>https:\/\/gloty-api.wp-shelf.com<\/code>, operated by WP Shelf. A developer can override it with <code>CITELY_SERVICE_URL<\/code> for a self-hosted service.<\/li>\n<li>Licensing requests send the license key, site URL, product identifier and, on activation, plugin version. They occur when you activate or deactivate a license and when cached license status is refreshed. These requests are separate from starting a rewrite.<\/li>\n<li>Starting a rewrite sends the original passage or heading, the page title, URL and post type, language, other headings, bounded body or section text, available audit signals, rewrite constraints, item identifiers and a batch identifier, together with the site URL and license key. A developer can change context through the <code>citely_rewrite_context<\/code> filter.<\/li>\n<li>The plugin sends batches after you start a run and polls for results while it is running. If enabled on the service, results may also arrive through an authenticated callback. Merely reading an audit or serving a page does not send its content for generation.<\/li>\n<li>The generation implementation uses OpenAI to process content. The plugin calls WP Shelf, not OpenAI directly. License keys are for WP Shelf authentication, not part of the provider's generation prompt.<\/li>\n<li>WP Shelf terms: https:\/\/wp-shelf.com\/terms<\/li>\n<li>WP Shelf privacy policy: https:\/\/wp-shelf.com\/privacy<\/li>\n<li>OpenAI terms: https:\/\/openai.com\/policies\/terms-of-use\/<\/li>\n<li>OpenAI privacy policy: https:\/\/openai.com\/policies\/privacy-policy\/<\/li>\n<\/ul>\n\n<p><strong>3. Optional usage reporting.<\/strong> Only after you enable usage reporting in\nSettings, a weekly background request sends the payload below to\n    https:\/\/gloty-api.wp-shelf.com\/v1\/telemetry (or your configured service host).\nIt is operated by WP Shelf under the same terms and privacy policy. Turning it\noff stops future reports. It is independent of license activation and rewriting.<\/p>\n\n<h4>What the optional report sends<\/h4>\n\n<p>No usage report is sent unless you enable it in Settings. Its JSON payload\ncontains the product identifier and the following data:<\/p>\n\n<ul>\n<li>How many AI crawlers you block, and how many of those are search crawlers \u2014\nas counts. Which crawlers you chose is your editorial decision and is never\nsent<\/li>\n<li>Whether WordPress's \"discourage search engines\" setting is on<\/li>\n<li>Whether you publish an llms.txt<\/li>\n<li>Once you have run an audit: how many pages were analysed, the median score,\nand the component counts behind it. A site that has never run one sends no\nscores at all rather than zeroes \u2014 an unmeasured site is not a badly\nstructured one<\/li>\n<li>Your WordPress, PHP and Crawlune versions, your site's locale, and whether\nWordPress considers the site production, staging or local<\/li>\n<li>A one-way hash identifying this installation, salted with a random value\ngenerated on your site and never sent. It lets one site's weekly reports be\ncounted as one site; it cannot be turned back into an address<\/li>\n<\/ul>\n\n<p>Not included in the usage-report payload: your site address, any e-mail address, any user account, any page,\nheading, answer or other content, any URL of yours, which specific crawlers you\nblock, and your license key. Turning the setting off stops the weekly job\nimmediately; uninstalling deletes the salt, so a later reinstall is a different\ninstallation as far as we can tell. As with any HTTPS request, the receiving\nserver can see your server's network address; the payload does not make the\nconnection anonymous.<\/p>\n\n<p><strong>E-mailed reports:<\/strong> the separate scheduled content report uses your site's\nWordPress mail configuration and the recipient you choose. Its delivery may use\nyour hosting or SMTP provider; Crawlune does not send that mail through WP Shelf.<\/p>\n\n<p>Why we ask: how many real sites arrive with a search crawler blocked decides\nwhether the crawler screen is the front door of this plugin or a footnote in it,\nand the median score in the wild is the only honest basis for saying what a\ntypical site's gap is. Both are currently guesses.<\/p>\n\n<p>Full documentation, including every component of the extractability score and what this plugin cannot measure: https:\/\/wp-shelf.com\/crawlune\/docs<\/p>\n\n<h4>Keeping your data when you uninstall<\/h4>\n\n<p>Deleting the plugin removes everything it stored, which is the right default \u2014 a\nplugin you have deleted should not leave rows behind. But it also removes your score history, which is a measurement series that cannot be recreated for months already past.<\/p>\n\n<p>If you are deleting the plugin to troubleshoot and intend to reinstall, add this\nto <code>wp-config.php<\/code> first:<\/p>\n\n<pre><code>define( 'CITELY_KEEP_DATA', true );\n<\/code><\/pre>\n\n<p>Uninstalling then leaves your data in place, and a reinstall picks it up where it\nleft off. Remove the line when you want a genuine clean removal.<\/p>\n\n<h3>Security<\/h3>\n\n<p>Found a security issue? Please email hello@wp-shelf.com rather than opening a\npublic support thread \u2014 a forum post is world-readable the moment you send it,\non a plugin installed on other people's sites. You will get an acknowledgement\nwithin 72 hours, and credit in the changelog unless you would rather not have it.<\/p>\n\n<h3>Source code<\/h3>\n\n<p>Readable JavaScript source is included in src\/ and shared-js\/, including the shared\nWP Shelf packages compiled into build\/. The ZIP includes package.json and\nwebpack.config.js. See shared-js\/README.md for rebuilding with Node.js and npm;\nno private repository or workspace checkout is required. WordPress supplies the\nReact and WordPress browser libraries used by the compiled admin interface.<\/p>\n\n<!--section=changelog-->\n<h4>0.62.6<\/h4>\n\n<ul>\n<li>Recover interrupted dispatches and partial result delivery without issuing duplicate paid work.<\/li>\n<li>Preserve saved results during a service-review pause and resume only on an explicit request.<\/li>\n<li>Continue recovering older multi-batch runs without requiring an open admin screen.<\/li>\n<li>Restore the complete rewrite history for each run, including large runs, while protecting newer rewrites.<\/li>\n<li>Preserve quotes and backslashes in saved article text and generated FAQ data.<\/li>\n<li>Report partial undo results accurately and keep unsuccessful restores available to retry.<\/li>\n<li>Show one-time allowances without a reset date, and display resets only when the service confirms renewal.<\/li>\n<\/ul>\n\n<h4>0.62.5<\/h4>\n\n<ul>\n<li>Escape assembled admin HTML at output while retaining form controls and nonces.<\/li>\n<li>Enforce non-HTML content types and nosniff for text and Markdown endpoints.<\/li>\n<li>Simplify JSON encoding flags while retaining JSON-LD script-injection protection.<\/li>\n<\/ul>\n\n<h4>0.62.4<\/h4>\n\n<ul>\n<li>Guide installations without a local license key to settings before creating a hosted rewrite run; existing runs and free tools remain available.<\/li>\n<li>Fixed: malformed service acknowledgements now use bounded retries instead of leaving a rewrite run waiting indefinitely.<\/li>\n<li>Fixed: unavailable results stop retrying after six hours even when the service reports the job complete; valid late results still arrive normally.<\/li>\n<li>Improved: failed batch messages explain the next step without implying that allowance was refunded.<\/li>\n<\/ul>\n\n<h4>0.62.3<\/h4>\n\n<ul>\n<li>Renamed the plugin to Crawlune \u2013 AI Crawler Controls, including its translation domain.<\/li>\n<li>Secured JSON-LD output against script-element breakout in page metadata and FAQ content.<\/li>\n<li>Added the submitting WordPress.org account to the contributor credits.<\/li>\n<\/ul>\n\n<h4>0.62.2<\/h4>\n\n<ul>\n<li>Fixed hosted generation and result polling on fresh installs with no saved server URL.<\/li>\n<li>Improved request-input sanitization and documented temporary CSV stream handling.<\/li>\n<li>Included shared JavaScript sources and standalone build instructions in the plugin ZIP.<\/li>\n<li>Clarified external-service requests and privacy disclosures.<\/li>\n<li>Verified clean installation and removal on WordPress 7.1.<\/li>\n<\/ul>\n\n<h4>0.62.0<\/h4>\n\n<ul>\n<li>Added JavaScript search, sorting and pagination for admin data tables.<\/li>\n<li>Navigate between plugin subpages asynchronously, with Back\/Forward support.<\/li>\n<\/ul>\n\n<p>Earlier release notes are included in CHANGELOG.md.<\/p>","raw_excerpt":"Control which AI crawlers may read your site, and audit whether your content can be quoted by answer engines.","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin\/366338","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin"}],"about":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/types\/plugin"}],"replies":[{"embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/comments?post=366338"}],"author":[{"embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wporg\/v1\/users\/wpshelf"}],"wp:attachment":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/media?parent=366338"}],"wp:term":[{"taxonomy":"plugin_section","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_section?post=366338"},{"taxonomy":"plugin_tags","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_tags?post=366338"},{"taxonomy":"plugin_category","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_category?post=366338"},{"taxonomy":"plugin_contributors","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_contributors?post=366338"},{"taxonomy":"plugin_business_model","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_business_model?post=366338"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}