{"id":366422,"date":"2026-09-30T07:41:48","date_gmt":"2026-09-30T07:41:48","guid":{"rendered":"https:\/\/wordpress.org\/plugins\/site-answers-llms-txt-ai-crawler-log-markdown\/"},"modified":"2026-09-30T08:42:14","modified_gmt":"2026-09-30T08:42:14","slug":"site-answers-llms-txt-ai-crawler-log-markdown","status":"publish","type":"plugin","link":"https:\/\/wordpress.org\/plugins\/site-answers-llms-txt-ai-crawler-log-markdown\/","author":23546237,"comment_status":"closed","ping_status":"closed","template":"","meta":{"version":"1.0.1","stable_tag":"1.0.1","tested":"7.1.2","requires":"6.5","requires_php":"7.4","requires_plugins":null,"header_name":"Site Answers \u2013 llms.txt, AI Crawler Log & Markdown","header_author":"lalokupfer","header_description":"Publishes \/llms.txt from your own published content, serves a clean Markdown copy of any page at the same address plus .md, and logs which AI crawlers visited and what they read.","assets_banners_color":"112035","last_updated":"2026-09-30 08:42:14","external_support_url":"","external_repository_url":"","donate_link":"","header_plugin_uri":"","header_author_uri":"https:\/\/profiles.wordpress.org\/lalokupfer\/","rating":0,"author_block_rating":0,"active_installs":0,"downloads":69,"num_ratings":0,"support_threads":0,"support_threads_resolved":0,"author_block_count":0,"sections":["description","installation","faq","changelog"],"tags":{"1.0.0":{"tag":"1.0.0","author":"lalokupfer","date":"2026-09-30 07:41:29","revision":3720571},"1.0.1":{"tag":"1.0.1","author":"lalokupfer","date":"2026-09-30 08:42:14","revision":3720719}},"upgrade_notice":{"1.0.1":"<p>Adds llms-full.txt, a robots.txt check for AI crawlers, and per-crawler Allow \/ Block choices. Your robots.txt does not change unless you choose.<\/p>","1.0.0":"<p>First release.<\/p>"},"ratings":[],"assets_icons":{"icon-128x128.png":{"filename":"icon-128x128.png","revision":3720570,"resolution":"128x128","location":"assets","locale":"","width":128,"height":128},"icon-256x256.png":{"filename":"icon-256x256.png","revision":3720570,"resolution":"256x256","location":"assets","locale":"","width":256,"height":256}},"assets_banners":{"banner-1544x500.png":{"filename":"banner-1544x500.png","revision":3720570,"resolution":"1544x500","location":"assets","locale":"","width":1544,"height":500},"banner-772x250.png":{"filename":"banner-772x250.png","revision":3720570,"resolution":"772x250","location":"assets","locale":"","width":772,"height":250}},"assets_blueprints":{},"all_blocks":[],"tagged_versions":["1.0.0","1.0.1"],"block_files":[],"assets_screenshots":{"screenshot-1.png":{"filename":"screenshot-1.png","revision":3720717,"resolution":"1","location":"assets","locale":"","width":1078,"height":1304},"screenshot-2.png":{"filename":"screenshot-2.png","revision":3720717,"resolution":"2","location":"assets","locale":"","width":1000,"height":430},"screenshot-3.png":{"filename":"screenshot-3.png","revision":3720717,"resolution":"3","location":"assets","locale":"","width":1078,"height":1823}},"screenshots":{"1":"The Overview: your llms.txt and llms-full.txt, whether your robots.txt lets AI crawlers in, which ones visited, and the pages they read most.","2":"The llms.txt file itself, built from your own posts and pages.","3":"Settings: which content is listed, the full-text file, and an Allow \/ Block \/ No change choice for each AI crawler."}},"plugin_section":[],"plugin_tags":[244526,232780,226124,244604,4608],"plugin_category":[],"plugin_contributors":[275437],"plugin_business_model":[],"class_list":["post-366422","plugin","type-plugin","status-publish","hentry","plugin_tags-aeo","plugin_tags-ai-crawler","plugin_tags-llm","plugin_tags-llms-txt","plugin_tags-markdown","plugin_contributors-lalokupfer","plugin_committers-lalokupfer"],"banners":{"banner":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/banner-772x250.png?rev=3720570","banner_2x":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/banner-1544x500.png?rev=3720570","banner_rtl":false,"banner_2x_rtl":false},"icons":{"svg":false,"icon":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/icon-128x128.png?rev=3720570","icon_2x":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/icon-256x256.png?rev=3720570","generated":false},"screenshots":[{"src":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/screenshot-1.png?rev=3720717","caption":"The Overview: your llms.txt and llms-full.txt, whether your robots.txt lets AI crawlers in, which ones visited, and the pages they read most."},{"src":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/screenshot-2.png?rev=3720717","caption":"The llms.txt file itself, built from your own posts and pages."},{"src":"https:\/\/ps.w.org\/site-answers-llms-txt-ai-crawler-log-markdown\/assets\/screenshot-3.png?rev=3720717","caption":"Settings: which content is listed, the full-text file, and an Allow \/ Block \/ No change choice for each AI crawler."}],"raw_content":"<!--section=description-->\n<p>AI assistants are reading your website. Site Answers has one job: make your site\nreadable to them, show you whether they actually came, and let you decide which of them\nmay. These parts do that job, and together they are the whole plugin.<\/p>\n\n<p><strong>It publishes an <code>llms.txt<\/code> file.<\/strong> <code>llms.txt<\/code> is an emerging convention \u2014 a plain,\nreadable Markdown index at the root of your site that tells AI assistants what you\nhave published and where to find it. Site Answers builds yours from your own posts and\npages: your site name, your tagline, then one line per URL with its title and a short\ndescription. It is rebuilt automatically whenever you publish, edit or delete\nsomething, so it never goes stale.<\/p>\n\n<p><strong>It publishes <code>llms-full.txt<\/code> too.<\/strong> Where <code>llms.txt<\/code> is the map, <code>llms-full.txt<\/code> is the\nwhole territory: the full text of every page <code>llms.txt<\/code> lists, one after another, in one file,\nso an AI assistant can read your site in a single visit. Same pages, same exclusions, capped at\n5 MB with a note saying where the rest is. You can switch it off in Settings.<\/p>\n\n<p><strong>It logs which AI crawlers visit.<\/strong> GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot,\nCCBot and the others are recognised by name. For each visit you see which crawler\ncame, which page it asked for, and whether it got the page or an error. That answers a\nquestion site owners currently cannot answer at all: <em>is AI search actually reading my\ncontent, and which pages does it want?<\/em><\/p>\n\n<p><strong>It serves a clean Markdown copy of every page.<\/strong> Add <code>.md<\/code> to the end of any post or page\naddress and you get the article as plain Markdown \u2014 headings, paragraphs, lists, links and\ntables \u2014 with the navigation, sidebars, share buttons and scripts left out. That is what an\nautomated reader wants, and it is far less likely to be misquoted than a page it had to dig out\nof a theme. Each HTML page also advertises its Markdown twin in a <code>Link<\/code> header, so a crawler\ncan find it without guessing.<\/p>\n\n<p><strong>It tells you when your robots.txt is turning AI crawlers away.<\/strong> An empty crawler log often\nhas a simple cause: a line in <code>robots.txt<\/code> \u2014 added by another plugin, a host, or years ago by\nhand \u2014 asking GPTBot or ClaudeBot to stay away. The Overview reads the robots.txt your site\nactually serves, the way the crawlers read it, and names every AI crawler it blocks.<\/p>\n\n<p><strong>And it lets you choose, crawler by crawler.<\/strong> In Settings, each AI crawler has three\nchoices: No change, Allow or Block. Block ChatGPT's training crawler and keep the one that\nfetches your page when someone asks ChatGPT about you, or the other way round. Your choices go\ninto WordPress's own robots.txt. Until you pick something, robots.txt stays exactly as it was.<\/p>\n\n<h4>How this is different<\/h4>\n\n<p>Plenty of plugins will write an <code>llms.txt<\/code> for you. Almost all of them stop at that point: the\nfile is written, and you never find out whether anything read it. This one closes the loop \u2014 it\npublishes the file, serves your pages in a form a machine can actually read, and then shows you\nwhether any AI crawler turned up and what it asked for. Publishing without measuring is\nguesswork, and until now measuring meant reading raw server logs.<\/p>\n\n<p>Three further differences, each of which you can check in the code:<\/p>\n\n<ul>\n<li><strong>It stands aside rather than fighting for the address.<\/strong> If another plugin, or a static file\nin your web root, already serves <code>\/llms.txt<\/code>, Site Answers never registers its own route for it:\ntheirs keeps working untouched, and your settings screen tells you which one is answering. Two\nplugins answering one address is how sites break, and it is always the newer plugin that gets\nblamed for it.<\/li>\n<li><strong>It keeps nothing personal.<\/strong> The crawler log holds the crawler's name, the page path, the\nresponse status and the time. No IP addresses, no query strings, no headers \u2014 so it never grows\ninto a consent-banner problem the way a general-purpose visitor log does.<\/li>\n<li><strong>It asks for nothing.<\/strong> No account, no API key, no external service, no outbound request of\nany kind, and nothing to upgrade to. Everything happens on your own server, so nothing about\nyour site or your visitors leaves it.<\/li>\n<\/ul>\n\n<h4>What it does not do<\/h4>\n\n<p>This plugin does not connect to any external service. It makes no outbound HTTP request\nof any kind \u2014 no API, no phone-home, no telemetry, no analytics, no third-party\nlibraries. Everything it shows you was produced by your own site, on your own server.<\/p>\n\n<p>It stores no IP addresses. The crawler log records the crawler's name, the page path and\nthe time \u2014 nothing else. Anything after the <code>?<\/code> in an address is discarded before the\nrecord is written, because query strings on real sites carry e-mail addresses, order\nreferences and reset tokens. Records older than 90 days are deleted automatically.<\/p>\n\n<p>It changes nothing on its own. Your robots.txt is left exactly as it is until you choose\nAllow or Block for a crawler, and no file is ever written to your server: if you have a\nrobots.txt file of your own, the plugin tells you so and leaves it alone.<\/p>\n\n<h4>Why this matters<\/h4>\n\n<p>AI assistants increasingly answer questions using content they read from websites, and being\nreadable to them is becoming its own discipline \u2014 answer engine optimization, or AEO, alongside\nthe generative engine optimization people now talk about next to ordinary SEO. Both come down\nto the same two practical questions: can a large language model find and parse what you\npublished, and is it actually doing so? Site Answers gives you a straight answer to the second\none, and does the groundwork for the first.<\/p>\n\n<h4>Who it is for<\/h4>\n\n<p>Anyone who publishes writing and wants to know whether AI search engines are reading it.\nBloggers, documentation sites, shops with real content, agencies who need to show a\nclient what is happening.<\/p>\n\n<h4>Privacy<\/h4>\n\n<p>No personal data of any kind is collected, stored or transmitted. Because no IP address\nis recorded and nothing leaves your server, this plugin does not put your site into\nconsent-banner territory.<\/p>\n\n<!--section=installation-->\n<ol>\n<li>Install and activate the plugin.<\/li>\n<li>Visit <strong>Site Answers \u2192 Overview<\/strong>. Your <code>llms.txt<\/code> link is at the top.<\/li>\n<li>Optionally visit <strong>Site Answers \u2192 Settings<\/strong> to choose which content is listed and which AI crawlers may read your site.<\/li>\n<\/ol>\n\n<p>There is nothing to sign up for and no key to enter.<\/p>\n\n<!--section=faq-->\n<dl>\n<dt id=\"what%20is%20llms.txt%3F\"><h3>What is llms.txt?<\/h3><\/dt>\n<dd><p>It is a proposed convention \u2014 like <code>robots.txt<\/code> or <code>sitemap.xml<\/code>, but written for large\nlanguage models. A single Markdown file at <code>\/llms.txt<\/code> that tells an AI assistant what a\nsite contains, in a form it can read without wading through navigation, adverts and\nscripts. It is young and not every AI engine reads it yet, which is exactly why it costs\nalmost nothing to publish one now.<\/p><\/dd>\n<dt id=\"where%20is%20my%20llms.txt%3F\"><h3>Where is my llms.txt?<\/h3><\/dt>\n<dd><p>At <code>\/llms.txt<\/code> on your site, and the Overview screen links straight to it.<\/p><\/dd>\n<dt id=\"what%20is%20llms-full.txt%3F\"><h3>What is llms-full.txt?<\/h3><\/dt>\n<dd><p>The companion to <code>llms.txt<\/code>. Instead of a list of links, it holds the full text of every page\n    llms.txt lists, as Markdown, in one file at <code>\/llms-full.txt<\/code>. Each page starts with its title\nand a <code>Source:<\/code> line with its address. It leaves out exactly what <code>llms.txt<\/code> leaves out \u2014\ndrafts, private and password-protected posts, and pages marked \"noindex\" \u2014 and stops at 5 MB.\nOn a large site the first visit may not include every page yet: pages are converted a few at a\ntime so the file never times out, and the file says so when that happens.<\/p><\/dd>\n<dt id=\"my%20llms.txt%20says%20%22not%20found%22.%20why%3F\"><h3>My llms.txt says \"not found\". Why?<\/h3><\/dt>\n<dd><p>Almost always because <strong>Settings \u2192 Reading \u2192 \"Discourage search engines from indexing\nthis site\"<\/strong> is switched on. Site Answers honours that: if you have asked search engines\nto stay away, it does not publish an index of your content on your behalf. The Overview\nscreen tells you when this is the reason. Switch that setting off and the file appears\nimmediately.<\/p>\n\n<p>If that setting is already off, re-save <strong>Settings \u2192 Permalinks<\/strong> once \u2014 that rebuilds\nWordPress's internal address table.<\/p><\/dd>\n<dt id=\"which%20ai%20crawlers%20are%20detected%3F\"><h3>Which AI crawlers are detected?<\/h3><\/dt>\n<dd><p>GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-User and anthropic-ai\n(Anthropic); PerplexityBot and Perplexity-User (Perplexity); Google-Extended and\nGoogleOther (Google); CCBot (Common Crawl); Bytespider (ByteDance); Amazonbot (Amazon);\nApplebot-Extended (Apple); Meta-ExternalAgent (Meta); cohere-ai (Cohere); DuckAssistBot\n(DuckDuckGo); YouBot (You.com); and Diffbot.<\/p>\n\n<p>Google-Extended and Applebot-Extended are robots.txt names only: Google and Apple fetch pages\nwith their ordinary crawlers and use these names to let you opt out of AI training. They can\nbe blocked in Settings, but they never appear in the visit log.<\/p><\/dd>\n<dt id=\"how%20do%20i%20get%20the%20markdown%20version%20of%20a%20page%3F\"><h3>How do I get the Markdown version of a page?<\/h3><\/dt>\n<dd><p>Add <code>.md<\/code> to the end of its address. A post at <code>https:\/\/example.com\/hello-world\/<\/code> is also\nserved at <code>https:\/\/example.com\/hello-world.md<\/code>. The HTML page carries a\n    Link: &lt;\u2026&gt;; rel=\"alternate\"; type=\"text\/markdown\" header pointing at it, which is how an\nautomated reader is meant to discover it.<\/p>\n\n<p>Drafts, private posts, password-protected posts and anything you have marked \"noindex\"\nare not served this way, exactly as they are kept out of <code>llms.txt<\/code>. The Markdown copy\nalso carries <code>X-Robots-Tag: noindex<\/code> and a canonical link back to the real page, so it\ncannot compete with your own article in search results.<\/p><\/dd>\n<dt id=\"does%20this%20plugin%20send%20anything%20to%20anyone%3F\"><h3>Does this plugin send anything to anyone?<\/h3><\/dt>\n<dd><p>No. It makes no outbound HTTP request at all. There is no account, no API key, no\ntelemetry and no analytics of any kind.<\/p><\/dd>\n<dt id=\"does%20it%20record%20ip%20addresses%3F\"><h3>Does it record IP addresses?<\/h3><\/dt>\n<dd><p>No. The crawler name, the page path and the time \u2014 that is the whole record. Query\nstrings are discarded too.<\/p><\/dd>\n<dt id=\"my%20server%20log%20shows%20more%20ai%20crawler%20visits%20than%20this%20plugin%20does.%20why%3F\"><h3>My server log shows more AI crawler visits than this plugin does. Why?<\/h3><\/dt>\n<dd><p>If your site uses full-page caching \u2014 nginx FastCGI cache, WP Rocket, Cloudflare, a host\nlevel cache \u2014 then a cached page is served without WordPress running at all, so the\nplugin never sees that visit. This is true of every WordPress-level bot log, not just\nthis one. Treat the numbers here as a floor rather than a total.<\/p><\/dd>\n<dt id=\"does%20it%20work%20with%20my%20seo%20plugin%3F\"><h3>Does it work with my SEO plugin?<\/h3><\/dt>\n<dd><p>Yes. If a page is marked \"noindex\" in Yoast SEO, Rank Math, SEOPress or The SEO\nFramework, Site Answers leaves it out of <code>llms.txt<\/code> by default. You can turn that off in\nSettings. For any other rule, the <code>siteanswers_llms_txt_exclude_post<\/code> filter lets you\nexclude a page in code.<\/p><\/dd>\n<dt id=\"will%20it%20list%20my%20drafts%20or%20private%20posts%3F\"><h3>Will it list my drafts or private posts?<\/h3><\/dt>\n<dd><p>No. Only published content. Drafts, pending, private, trashed and password-protected\nposts are never listed, whatever you choose in Settings.<\/p><\/dd>\n<dt id=\"what%20happens%20when%20i%20delete%20the%20plugin%3F\"><h3>What happens when I delete the plugin?<\/h3><\/dt>\n<dd><p>Everything it created is removed: its database table, every option and cached value it\nwrote, and its scheduled cleanup job. Deactivating the plugin, by contrast, leaves your\ncrawler history intact so you can switch it back on without losing anything.<\/p><\/dd>\n<dt id=\"can%20i%20block%20ai%20crawlers%20like%20gptbot%3F\"><h3>Can I block AI crawlers like GPTBot?<\/h3><\/dt>\n<dd><p>Yes, if you choose to. <strong>Site Answers \u2192 Settings<\/strong> has a choice for every crawler it knows:\nNo change, Allow or Block. Block adds a <code>Disallow: \/<\/code> rule for that crawler to your\nrobots.txt; Allow adds a rule inviting it in. Nothing changes until you choose, and each\ncrawler's row says what it is for \u2014 some collect training data, others open your page only\nwhen someone asks an AI assistant about it, and blocking those means the assistant cannot read\nyou when asked.<\/p>\n\n<p>robots.txt is a request that well-behaved crawlers follow, not a lock. If your site has a\nrobots.txt file of its own on the server, WordPress's robots.txt is never used, so the choices\ncannot apply \u2014 the Settings screen tells you when that is the case.<\/p><\/dd>\n<dt id=\"why%20does%20the%20overview%20say%20my%20robots.txt%20blocks%20a%20crawler%20i%20never%20blocked%3F\"><h3>Why does the Overview say my robots.txt blocks a crawler I never blocked?<\/h3><\/dt>\n<dd><p>Because something else did: another plugin, your host, or a line in a robots.txt file on your\nserver. The Overview reads the robots.txt your visitors actually get and names every AI\ncrawler it turns away, whoever wrote the rule. If you set a crawler to Allow and another rule\nstill blocks it, the Overview says that too.<\/p><\/dd>\n\n<\/dl>\n\n<!--section=changelog-->\n<h4>1.0.1<\/h4>\n\n<ul>\n<li>New: <code>\/llms-full.txt<\/code> \u2014 the full text of every page in <code>llms.txt<\/code>, in one file. Same pages and exclusions, 5 MB cap, can be switched off.<\/li>\n<li>New: the Overview says when your robots.txt asks an AI crawler to stay away, whoever wrote the rule.<\/li>\n<li>New: allow or block each AI crawler in robots.txt from Settings. Nothing changes until you choose, and no file is written.<\/li>\n<li><code>llms.txt<\/code> now points to <code>llms-full.txt<\/code> while it is published.<\/li>\n<li>Tested on every PHP version from 7.4 to 8.4 with WordPress 6.5, 6.8, 7.0 and 7.1.<\/li>\n<\/ul>\n\n<h4>1.0.0<\/h4>\n\n<ul>\n<li>First release.<\/li>\n<li>Publishes <code>\/llms.txt<\/code> from your published posts and pages, rebuilt automatically on publish, edit and delete.<\/li>\n<li>Serves a clean Markdown copy of any post or page at the same address plus <code>.md<\/code>, and advertises it in a <code>Link<\/code> header.<\/li>\n<li>Logs visits from 19 named AI crawlers, with the page requested and the response status.<\/li>\n<li>Settings for content types, how many entries to list, and whether to honour \"noindex\".<\/li>\n<li>No external requests, no IP addresses, no telemetry.<\/li>\n<\/ul>","raw_excerpt":"Publish llms.txt and llms-full.txt, serve clean Markdown, and see which AI crawlers read your site. No account, no external service.","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin\/366422","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin"}],"about":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/types\/plugin"}],"replies":[{"embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/comments?post=366422"}],"author":[{"embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wporg\/v1\/users\/lalokupfer"}],"wp:attachment":[{"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/media?parent=366422"}],"wp:term":[{"taxonomy":"plugin_section","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_section?post=366422"},{"taxonomy":"plugin_tags","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_tags?post=366422"},{"taxonomy":"plugin_category","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_category?post=366422"},{"taxonomy":"plugin_contributors","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_contributors?post=366422"},{"taxonomy":"plugin_business_model","embeddable":true,"href":"https:\/\/wordpress.org\/plugins\/wp-json\/wp\/v2\/plugin_business_model?post=366422"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}