We sell an SEO and AI visibility agent. Until this month, when a site had no llms.txt, our agent would write one and open a pull request to add it. In May it did exactly that on our own site, and we merged it. So we had a reason to want the file to matter.
Then we read our own server logs. Over 98 days, the crawlers run by Anthropic and OpenAI visited our server 2,028 times. They asked for robots.txt 827 times. They asked for /llms.txt zero times, and for /llms-full.txt zero times.
This post covers how we counted, what could make the number wrong, and what we changed.
What llms.txt is
llms.txt is a proposal by Jeremy Howard, published on 3 September 2024 at llmstxt.org, "to standardise on using an /llms.txt file to provide information to help agents use a website." It is a Markdown file at the root of a site. Whether AI systems actually fetch it is a separate question, and it is the one we could test.
The data
We host aseoka.com and its subdomains on one server behind Cloudflare. Everything goes through one nginx instance, which writes one access log. We read that log from its first line to the moment we ran the count:
- Window: 2026-06-22 04:56:02 UTC to 2026-09-27 21:21:50 UTC (98 calendar days).
- Size: 991,125 requests, after we removed 8 requests we made ourselves while preparing this post.
- The file:
/llms.txtreturned HTTP 200 for the whole window. It had been live since 15 May 2026. Ourrobots.txtallowed Anthropic's, OpenAI's and Perplexity's crawlers to fetch it for the whole window.
How we told real AI crawlers from impostors
Anyone can put "GPTBot" or "ClaudeBot" in a user agent string, and many clients do. In our log, 21,099 requests claimed to come from an Anthropic, OpenAI, Perplexity or Common Crawl crawler. Only 2,032 of them came from IP addresses those companies publish for their crawlers. If you count by user agent, you mostly count impostors.
So we ignored the user agent for verification and used the IP address instead:
- Because the site sits behind Cloudflare, nginx sees a Cloudflare address as the client. We took the real client address from the
X-Forwarded-Forheader that Cloudflare adds. - We downloaded each company's published crawler IP list on the day we ran the count: Anthropic (which Anthropic describes as the list for ClaudeBot, Claude-User and Claude-SearchBot), OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Perplexity (user fetcher) and Common Crawl. For comparison we also used Google and Bing.
- A request counted as a real AI crawler only if its address was inside one of those lists.
- We separated out our own traffic: our agent checks whether
/llms.txtexists, so it requests the file from our own server.
Results
What verified crawlers requested, by the published IP list they came from:
| Crawler (verified by IP list) | All requests | robots.txt | sitemap.xml | /llms.txt | /llms-full.txt |
|---|---|---|---|---|---|
| Anthropic | 1,220 | 578 | 579 | 0 | 0 |
| OpenAI | 808 | 249 | 100 | 0 | 0 |
| Common Crawl | 3 | 3 | 0 | 0 | 0 |
| Perplexity | 1 | 1 | 0 | 0 | 0 |
| Google (search crawlers) | 581 | 163 | 0 | 0 | 0 |
| Bing | 280 | 46 | 65 | 0 | 0 |
These crawlers were not absent. Anthropic's crawler showed up in every month of the window (136 requests in the last nine days of June, 617 in July, 144 in August, 323 in September). OpenAI's did too (71, 203, 264, 270). They read robots.txt, the sitemap and our pages. They did not read /llms.txt.
Who did request /llms.txt (105 requests in total):
| Requester | Requests |
|---|---|
| Our own agent and link checker | 69 |
| Generic browser user agents from unlisted addresses | 21 |
| SEOJuice's crawler | 8 |
| curl from one home connection (almost certainly us, testing) | 3 |
| Dataprovider.com | 3 |
| PipericBot | 1 |
| Any request with an AI company's user agent, real or fake | 0 |
| Any request from an AI company's published IP range | 0 |
Two thirds of the attention our llms.txt got came from our own tool checking that it existed. The rest came from SEO tools and unidentified clients. Nobody even bothered to fake an AI crawler to fetch it.
"Maybe Cloudflare blocked them before they reached you"
This is the strongest objection, so here is what we can and cannot rule out.
What rules it out as a general explanation. The same Anthropic and OpenAI crawlers reached our origin server 2,028 times during the window, in every month, for robots.txt, the sitemap and ordinary pages on the same site. So Cloudflare was not blocking AI crawlers across the board. Caching is also unlikely to hide the requests: .txt is not on Cloudflare's list of file types it caches by default, and other clients' requests for /llms.txt kept reaching our server, including two 32 seconds apart.
What it cannot rule out. A Cloudflare rule that applies only to the /llms.txt path, or a challenge served at the edge only for that path, would never show up in our logs. Nothing about the other traffic suggests one, but our origin log cannot prove a negative about Cloudflare.
Limitations
- One site. This is our own marketing site, and it gets very little search traffic. A popular documentation site could see different behaviour.
- User-triggered fetchers. ChatGPT-User, Claude-User and Perplexity-User fetch a page when a person asks the assistant about it. On a small site that rarely happens, so their absence tells you little. The fair test is the crawlers that visit on their own schedule, and they came hundreds of times.
- Today's IP lists, applied to past months. If a company used an address in July that is no longer on its list, we would have missed it. That cannot change the
/llms.txtresult, though: none of the 105 requests carried an AI company's user agent at all. - One log for several subdomains. Our log format does not record the hostname. Every
/llms.txtrequest that got a 200 hit aseoka.com, because only that site serves the file. - Our file is not a textbook example. It is a plain list of URLs with comment headings, not the linked Markdown the proposal describes. That does not affect the result: no AI crawler requested the file, so none of them ever saw its contents.
- This does not show that no AI system anywhere uses llms.txt. It shows that on our site, over 98 days, the major AI crawlers did not request it.
What we changed
Our agent treats a missing llms.txt as report-only. It still tells you the file is missing, but it will not open a pull request to add one. We made that change on 24 September 2026 and it has been in production since our 25 September release. The line in our code that does it carries the comment "no major AI vendor commits to reading llms.txt".
To be precise about the order: we made that call on 24 September after reviewing the published evidence and our own pull request history. We ran this log study the next day, and it is why we are comfortable keeping the change. llms.txt has also never counted toward the AI visibility score our product reports, and it does not today.
What to do with your own site
- Don't expect llms.txt to change your AI visibility. It is cheap to keep one if you have it, but on our evidence it is not something to prioritise.
- Look at your own logs before you believe any chart. Check what real crawlers ask for. On our site that meant
robots.txt, the sitemap and pages. - Verify crawlers by IP address, not user agent. Most requests in our log that claimed to be an AI crawler were not.
- Spend the effort where crawlers actually go. Make sure
robots.txtallows the AI search crawlers you want, keep your sitemap accurate, and make sure important content is in the HTML the server sends, not only rendered in the browser.
If you want to see how AI crawlers and assistants see your site, our free AI Visibility Audit checks what they can reach on a page and what they can't. Enter a URL and you get a summary in about 20 seconds; the full report is unlocked with your email address.