Every guide to AI search, ours included, tells you to let the crawlers in. Almost none of them tell you what happens after you do, and the reason is boring: nobody keeps the logs.
Cloudflare's free plan holds eight days of bot analytics and caps a single query at a one-day window. That is enough to answer "did GPTBot come by this week" and not enough to answer anything larger. So any claim about how AI crawlers behave over a quarter is either built on somebody's own daily snapshots or it is built on nothing. The data cannot be backfilled later, which is the part that makes it worth publishing: you either started saving it or you did not.
We started saving it. ConceptSEO writes a per-crawler request count to its own database once a day, for every domain it manages, read from the Cloudflare edge and scoped to that domain's hostname. This page reports 90 days of that record. Every number below came out of a query against it, the sample size sits next to each one, and where the record cannot answer something the question is left open rather than filled in.
Requests by crawler over the window. Figures as tabulated below.
Where does Googlebot rank?
Fifth. Four crawlers fetched these sites harder than Google's did over the same 90 days.
| Crawler | Type | Requests | Sites seen on |
|---|---|---|---|
| YandexBot | Search engine | 284,725 | 16 of 16 |
| ChatGPT-User | AI, live | 228,110 | 16 of 16 |
| bingbot | Search engine | 108,221 | 16 of 16 |
| Meta-ExternalAgent | AI, training or indexing | 78,826 | 16 of 16 |
| Googlebot | Search engine | 70,343 | 16 of 16 |
| Bytespider | AI, training or indexing | 59,367 | 15 of 16 |
| Applebot | Search engine | 57,517 | 16 of 16 |
| ClaudeBot | AI, training or indexing | 33,441 | 16 of 16 |
| GoogleOther | AI, training or indexing | 31,360 | 16 of 16 |
| PerplexityBot | AI, training or indexing | 28,695 | 16 of 16 |
The busiest AI operator is ChatGPT-User, at 228,110 requests. That is 3.2 times what Googlebot fetched from the same sites across the same window. ChatGPT-User is not a training crawler. It is the fetch OpenAI makes when somebody in a chat window asks a question and the model goes to look at a page.
The single busiest crawler of all is not an AI operator at all. YandexBot made 284,725 requests, more than Googlebot and bingbot put together. We have no explanation for why Yandex crawls this portfolio as hard as it does, and rather than invent one, the number is simply here. If you have edge logs of your own, it is worth checking whether you see the same thing.
Worth naming what this table is not. These are request counts, not visits, sessions, or traffic. A crawler fetching a page 200 times is a fact about the crawler, not evidence that anyone read anything.
How much of this is somebody waiting for an answer?
The distinction that matters commercially is not which company owns the crawler. It is whether a human is currently sitting there. A training crawler is stocking a warehouse. A live fetch means someone asked a question about ninety seconds ago and the assistant went out to read your page before answering.
That is not the shape we expected, and it changes what the number is good for. It is strong evidence that people are asking assistants about the subjects these sites cover. It is weak evidence about AI search as a diversified channel, because 90 days of data here is really 90 days of ChatGPT with a rounding error attached.
Is one big site carrying this?
No, and this is the part we found most convincing. The ratio is not a portfolio average that a single heavy domain could drag upward. Every site in the set shows it independently.
On 16 of 16 sites, AI crawlers made more requests than Googlebot did. The median site ratio is 8.3 to 1. The lowest is 1.8 to 1 and the highest is 30.2 to 1, though that top figure comes from a site with only 12 days in the window and should be read as thin rather than extreme.
| Site | Days measured | AI, live | AI, training | Googlebot | AI per Googlebot request |
|---|---|---|---|---|---|
| Site A | 90 | 170,652 | 95,950 | 23,029 | 11.6 |
| Site B | 90 | 31,289 | 17,716 | 4,274 | 11.5 |
| Site C | 90 | 6,072 | 23,815 | 7,616 | 3.9 |
| Site D | 90 | 6,292 | 59,878 | 18,183 | 3.6 |
| Site E | 78 | 4,075 | 7,830 | 3,109 | 3.8 |
| Site F | 65 | 5,770 | 8,762 | 2,784 | 5.2 |
| Site G | 65 | 656 | 5,235 | 1,413 | 4.2 |
| Site H | 65 | 370 | 1,706 | 1,133 | 1.8 |
| Site I | 40 | 722 | 2,189 | 460 | 6.3 |
| Site J | 37 | 3,411 | 9,952 | 1,448 | 9.2 |
| Site K | 37 | 6,122 | 7,391 | 1,378 | 9.8 |
| Site L | 37 | 4,399 | 7,329 | 1,114 | 10.5 |
| Site M | 36 | 1,765 | 5,971 | 1,055 | 7.3 |
| Site N | 36 | 793 | 13,206 | 1,461 | 9.6 |
| Site O | 30 | 3,699 | 7,425 | 1,037 | 10.7 |
| Site P | 12 | 1,288 | 24,319 | 849 | 30.2 |
The live and training columns move independently, which is the other thing worth noticing. Site D takes almost ten times more training crawl than live fetching; Site A and Site B are the reverse. Whatever decides how often an assistant goes out to read a given site, it is not the same thing that decides how often that site gets harvested for training.
Which crawlers are getting turned away?
The record also stores how each crawler's requests were answered, and a 403 is the interesting one because it means the site refused. A 404 or a 410 is not a refusal, it is the site correctly saying a URL is gone, so both are excluded here. This section covers 9 August onward, because that is when the collector started splitting 403 out of the general error bucket, and it only lists crawlers with at least 20 requests in that period.
| Crawler | Requests | Refused (403) | Refusal rate |
|---|---|---|---|
| FacebookBot | 2,473 | 1,071 | 43.3% |
| Bytespider | 32,439 | 9,222 | 28.4% |
| Google-Extended | 1,967 | 534 | 27.2% |
| CCBot | 1,767 | 470 | 26.6% |
| PerplexityBot | 15,934 | 3,566 | 22.4% |
| Amazonbot | 14,902 | 2,617 | 17.6% |
| Claude-SearchBot | 376 | 60 | 16.0% |
| GPTBot | 5,580 | 676 | 12.1% |
| Perplexity-User | 866 | 91 | 10.5% |
| ChatGPT-User | 86,798 | 4,326 | 5.0% |
| OAI-SearchBot | 14,382 | 668 | 4.6% |
| ClaudeBot | 14,413 | 635 | 4.4% |
| Claude-User | 13,926 | 137 | 1.0% |
| MistralAI-User | 922 | 8 | 0.9% |
| DuckAssistBot | 2,504 | 2 | 0.1% |
| Meta-ExternalAgent | 75,649 | 51 | 0.1% |
| GoogleOther | 21,508 | 0 | 0% |
| anthropic-ai | 6,217 | 0 | 0% |
The spread runs from 0% to 43.3% on sites that are all configured by the same team, which was not what we expected to find. The obvious hypothesis is that refusals track intent, with training crawlers blocked and live assistant fetches allowed. The table does not support it. Meta-ExternalAgent and GoogleOther are both training crawlers and both sit at or near zero, while GPTBot, also a training crawler, is refused 12.1% of the time. Bytespider and PerplexityBot are refused at roughly 20 to 30% while ClaudeBot, doing the same job, is refused at 4.4%.
The likeliest explanation is per-zone configuration rather than anything about the crawler: managed bot rules, an AI-crawler toggle left on when a zone was created, a country rule catching one crawler's egress and not another's. We have not traced each refusal back to the specific rule that produced it, so this is a measurement and not a diagnosis. What it is good for is the sanity check it prompted here, which is worth doing on your own sites: if you believe you are open to AI crawlers, this is the number that tells you whether you actually are. Ours said we were not, in places.
What this does not say
The sample is 16 domains that one agency manages. It is a portfolio, not a sample of the web, and the mix is small business and niche publishing rather than news, ecommerce, or anything at scale. A crawler's behavior on a 30-page site is not evidence of its behavior on a 300,000-page one.
Coverage inside the window is uneven. Four sites have all 90 days; the rest were added during the window, which is why the totals rest on 898 site-days rather than the 1,440 that 16 complete sites would give. Each site's day count is printed in its row, and every cross-crawler comparison uses the same days for both sides.
Two mapping caveats, both of which push against the headline rather than for it. Records written before 7 August identify a crawler by display label rather than by signature, and in that era two search signatures were folded into their parents: Googlebot-Image counted as Googlebot, which inflates the Googlebot column, and Applebot-Extended counted as Applebot, which leaves a small amount of AI traffic outside the AI total. Correcting both would widen the 7.8 to 1 gap, not narrow it.
Finally, none of this measures traffic. Cloudflare's referrer dataset, the one that would show which visitors an assistant actually sent, needs a paid plan and is not collected here. Crawling is the thing being counted. Whether it converts into visits is a separate question this record cannot answer, and we would rather say so than let a large number stand in for one we do not have.
Republish the table
Here is the headline table as HTML, source line attached. It is the same array the visible table renders from, so the two cannot drift apart.
Sample. 16 domains under management, 898 measured site-days, 546,049 AI-crawler requests and 70,343 Googlebot requests.
Window. 10 June 2026 to 7 September 2026, 90 days. Every one of the 90 dates carries at least one site's data.
Measurement source. Cloudflare bot analytics, read once a day per domain and written to ConceptSEO's own database. Each daily record covers a single day of traffic and is scoped to that domain's hostname, so a zone serving several subdomains never pools them. Cloudflare itself retains eight days on the plans these zones use, so this history exists only because it was saved as it went and cannot be reconstructed after the fact.
Construction. Crawlers are grouped by the signature the collector matched, longest match first. Live means an operator that fetches in response to a person's question; training or indexing means everything else. The two are never added together except where the text says "AI crawlers" and gives both components. Refusal rates use 403 responses only, cover 9 August onward, and exclude any crawler under 20 requests in that period.
Anonymization. Sites are labeled A through P. The portfolio mixes client domains with properties this business owns, and rather than name some and not others, none are named and no site is ranked against another on anything but its own crawl mix.
Last updated. 8 September 2026. The window is open and these figures will move. The date above is the version you are reading.