Blog

The AI Crawlers Showed Up. Here's Who Actually Reads My Website.

Seven weeks ago I wrote that I'd made this website agent-ready — and I was honest about the part that stung: almost nothing was reading any of it yet. That post ended with a promise instead of a prediction. Every AI crawler that touches this site gets logged — which engine, which page, which file — and when something changed, I'd publish the receipts, page by page, engine by engine.

Something changed.

How the counting works — and its limits

Every request from an AI crawler gets logged at the edge — which engine it belongs to, what it asked for, and what kind of thing that was: a real page, a robots.txt or sitemap check, one of the plain-text files I publish specifically for AI, or a probe for something that was never on this site at all. More on that last category later, because it's the fun one.

No paid tools, no analytics suite. It runs on the same infrastructure that serves the site, and it costs nothing.

The limits, stated up front so you can weigh the rest: this is one site, a 250-page local web shop, over one 28-day window with seven weeks of history behind it — a case log, not a study of the whole internet. And a request declares whose bot it is the way a phone call declares the caller's name: it says whatever it wants. Some companies publish the network addresses their crawlers really come from, so their visits can be checked. Some don't. Keep that in your pocket, because somebody in these logs is wearing an AI crawler's name tag and is very much not an AI crawler.

Six weeks of almost nothing

When I checked the logs in mid-July, about a week into counting, the honest summary was: nobody's home.

The plain-text twins — a clean text copy of every page, built so AI doesn't have to wade through design and code — had zero real fetches. The only hits on them were my own checks from the day I set them up. Anthropic's crawler was the busiest name in the log, and about 94% of its visits were it re-reading robots.txt and my sitemaps — knocking on the door over and over without ever coming inside. OpenAI was the only company actually reading pages. And of the 119 city pages I'd built by then for the service-area section, exactly one had ever been visited by an AI crawler.

I had started drafting the sentence in my head: nice idea, no readers, keep it because it's free. That's the sentence this post was supposed to be.

Then, in about a month, all of it flipped

The last 28 days on this site: 4,349 AI-crawler hits, and 1,971 of them were full page reads.

ClaudeBot went from knocking on the door to reading more or less the whole house — 1,215 page reads on its own, more than every other AI crawler here combined. It worked through essentially the entire location section: 379 distinct addresses under /areas/, which is more pages than that section even has today, because it's still dutifully re-checking ones I retired in August. OpenAI splits its traffic across three bots — one indexing for training, one feeding its search engine, one fetching a page live because a real person's chat asked about it — and all three were active here. Perplexity, absent from July's logs entirely, made 400 visits. Even Google's AI-side crawlers — separate from the regular Googlebot that's indexed this site all along — showed up in small numbers.

Nothing about this site got popular in those weeks. Human traffic here is still measured in dozens, and I publish that fact on purpose. What changed is which machines consider a small local business site worth reading. The full engine-by-engine count is in the table below.

The numbers, if you're citing this

Data posts get quoted, so here's the digest, and the terms: every figure on this page is free to reuse, commercial or not — just link back to this post as the source. If your own logs say something different, publish them and tell me; I'll link the counter-example from here.

One small-business website, 250 pages, 28 days ending late August 2026. Total AI-crawler hits: 4,349 — of which 1,971 were full page reads, 1,680 were robots.txt, sitemap and other housekeeping checks, roughly 600 were fake AI bots probing for secrets, and 96 were fetches of the llms.txt-style plain-text files. That llms.txt figure was zero for the first six weeks after launch. ClaudeBot alone read 1,215 pages — 62% of all AI page reads here. Live user-triggered fetches, where a real person's AI assistant pulled a page mid-conversation: 694 logged across ChatGPT and Claude, precise total soft because impostors wear those names too. Confirmed human click-throughs from chatgpt.com: two — who stayed an average of four and a half minutes against a ten-second site average, roughly 27 times longer. Small numbers, stated as small. That's the point of publishing them.

llms.txt, measured: six weeks of zero, then 86 fetches

Alongside every normal page, this site publishes a plain-text twin, plus an index file of all of them — the llms.txt convention, if you've run into that term. It's the piece of AI-readiness the internet argues about most, and almost everything written about it is opinion, in both directions. Here's a log instead.

First six weeks: zero fetches. Last 28 days: 86 fetches of the twins and roughly ten of the index files. OpenAI leads with 41, Anthropic right behind at 39, and the rest scatter across Meta, Perplexity — and, in a couple of cases, crawlers wearing Google's badge, which deserves a careful sentence: Google has said it doesn't use these files, and a fetch doesn't contradict that. Fetching isn't using.

That cuts both ways, so I'll say it plainly: I can prove the twins get read now. I cannot prove a single citation came from them. The reason I keep them anyway hasn't changed since July — they're generated automatically from the same pages the humans see, so they cost nothing to maintain and can't drift out of date. If someone tries to sell you llms.txt as magic, ask for their logs. Mine say: not magic, no longer dead.

The part I actually care about: the humans

Two of the bots in these logs aren't really crawlers. ChatGPT-User and Claude-User fire when an actual person, mid-conversation, asks something that makes the AI go fetch a page right then. Every one of those hits is a human being whose assistant just read part of my website to them.

The logs show 504 of those fetches under ChatGPT's live-user agent and 190 under Claude's in 28 days. I won't pretend every one is real — the fakers in the next section wear these exact name tags, so treat the precise totals loosely. But the clean reads are unambiguous, because of what they asked for: the homepage, the remote-work page, the local SEO service page, my work portfolio, one specific city page, and — my favorite — two specific blog posts, including the nichest thing I've ever written, a pricing guide for court-reporting websites. Somebody asked their AI a question, and the AI went and read that.

And on the other side of the counter, my analytics caught visitors arriving from chatgpt.com itself: two of them last month. Two is a number I'd be embarrassed to put on a slide, which is exactly why it's here — but those two stayed four and a half minutes each on a site where the average visit is ten seconds. Twenty-seven times the engagement, population of two. The trickle is real, the trickle is tiny, and it's the most interested traffic this site gets.

About one hit in seven was a fake AI bot

Roughly 600 of the 4,349 hits — about one in seven — weren't reading anything. They were probing. Requests for /.git/config, for environment files with passwords in them, even for /etc/passwd through an old development-server trick. Somebody's vulnerability scanner rotates through the names of half a dozen AI crawlers — ClaudeBot, CCBot, ChatGPT's live-user agent — hoping the AI costume keeps it off blocklists while it hunts for leaked credentials.

On this site every one of those requests hits a static wall and gets a 404 — there's no server to break into and no secrets at guessable addresses. But the lesson generalizes, and it's the thing I'd want quoted from this section: user-agent strings are self-reported, so a chunk of what gets sold as "AI traffic" is a scanner in a costume. The tells are the paths — no real reader asks for your .git folder — and the network addresses, which the real OpenAI, Google, and Perplexity publish and impostors can't use.

So when a tool or a consultant announces that AI traffic to your site is up 40%, the correct next question is: verified how?

What I'd actually do about this if I ran a small business

Not much — but not nothing, and the order matters.

Don't block AI crawlers by default. I know the instinct. But people are already asking these assistants who to hire, and an assistant can only recommend what it was allowed to read. My robots.txt allows every major AI reader by name, on purpose, and this log is what that policy earns: the engines people actually ask are reading the pages I'd want quoted.

Make the site readable the boring way. Static pages, fast, real text — every claim visible in the HTML instead of locked behind scripts. The engines that read 1,971 pages here didn't run my JavaScript to do it. If your site only works with scripts on, to an AI reader it barely exists. Put straight answers on the page — what you do, what it costs, where you work — because that's what gets read back to the person asking.

Then let it run quietly and count. This is the same discipline as local SEO — being legible to the systems people use to choose a business — extended to one more system. The loop is real: crawler reads page, person asks assistant, assistant fetches page, person shows up. I've now watched every step happen in my own logs. It's a trickle, and I won't dress it up as more. But every earlier shift looked exactly like this at the start, and the plumbing that catches it costs nothing to keep on.

Everything this site publishes for AI — the twins, the structured answers, the tools, with links to verify each claim — is written up at the agent-readiness page. If you want the same count running on your own site, or a site built to be read by this second kind of reader in the first place, that's a conversation I'm happy to have — plain terms, real numbers, same as always.

EngineHits (28 days)Full page readsWhat changed since July
Anthropic — ClaudeBot + Claude-User2,3731,215Was ~94% robots.txt polling; now reads nearly everything
OpenAI — GPTBot, OAI-SearchBot, ChatGPT-User1,368587Only reader in July; now also fetching live for users
Perplexity — PerplexityBot40094Didn't exist in July's logs
Google AI — Google-Extended, GoogleOther18457Trace in July; still small
Meta, Common Crawl, ByteDance2418Trace amounts

Frequently asked questions

Do AI crawlers actually visit small business websites?

Yes. This site — a one-person local web shop, about 250 pages — logged 4,349 AI-crawler hits in 28 days, and 1,971 of them were full page reads. That's with human traffic measured in dozens. The reading is real; it's the dollar value that's still small.

Should I block AI crawlers from my website?

My take: no, unless you have a specific reason. People already ask AI assistants who to hire, and an assistant can only recommend what it was allowed to read. Blocking makes you invisible to that channel to protect content that, for most local businesses, exists precisely to be found.

Does llms.txt actually do anything?

Honest answer from my own logs: the files went from zero fetches in six weeks to 96 in a month, led by OpenAI and Anthropic — so they are read. I can't prove a single citation resulted. Mine cost nothing because they're generated from the normal pages automatically. Free and read is worth keeping; I wouldn't pay much for magic.

Is ChatGPT sending real customers to websites?

Real visitors, yes — a trickle. My analytics caught two arrivals from chatgpt.com last month, and they averaged four and a half minutes on a site where the typical visit is ten seconds — about 27 times the engagement, on a population of two. Anyone claiming big AI-referral numbers for a small business should show you logs, not vibes.

Can I cite the numbers in this post?

Yes — any figure here, commercial use included, with a link back to this post as the source. The method and limits are described above: one site, 28 days, user-agent-based counting with impostors filtered by path and, where published, network address. If your logs disagree, publish them and I'll link yours too.