Fake AI Crawlers: They Are Hunting Your Credentials, Not Your Content

Fake AI crawlers are requests that wear the name of a trusted crawler such as Googlebot, Bytespider, or GPTBot without coming from the operator behind that name. Across the sites citAEOtion monitors, 23,614 crawler requests failed verification as of September 2026, and one in three of those forgeries never read a page. They went straight for credential files.
Quick answer: A user-agent is a text field the sender fills in, so any request can call itself Googlebot or Bytespider. When every claimed crawler identity on our network is checked, roughly 1 in 14 turns out to be a forgery. Those fakes are not training a model. 33 percent of them probed for passwords, API keys, and environment files, and across 17,007 credential scans wearing a trusted crawler's name, not one came from a verified operator. The real crawlers read your pages. The fakes try your doors.
Everyone Watches the Real Crawlers. Nobody Watches the Fakes.
If you care about AI search, you already ask which crawlers read your content and whether the answer engines cite you back. That is the right question. It has a hidden assumption: that the GPTBot, Googlebot, and Bytespider entries in your logs are who they say they are.
A large share are not. And the ones that are not behave nothing like a content crawler.
citAEOtion tracks AI and search crawler traffic across a fleet of live sites. Every request that claims a crawler identity is tested and returned with a verdict: verified, impostor, or unverifiable. As of September 2026 the count stands at 377,387 crawler requests observed, 317,424 verified as genuine, 23,614 caught as impostors, and 35,251 that claimed a name they could not back up. The fakes are the subject of this page.
A Third of the Fake AI Crawlers Went for Your Secrets
Of every request that failed verification, 33 percent never touched an article, a product page, or a post. They requested credential files. Environment files. Cloud metadata endpoints. Configuration and deploy secrets. The files that hold database passwords and API keys.
Here is what that traffic was reaching for, all time, across the sites we monitor:
| What the request was after | Requests | From proven impostors | From unverifiable claims |
|---|---|---|---|
| Credential and secrets files | 17,007 | 7,463 | 8,575 |
| Cloud metadata (SSRF) | 883 | 263 | 588 |
| Admin and login endpoints | 287 | 33 | 21 |
| Path traversal | 135 | 34 | 86 |
| Code execution attempts | 7 | 0 | 1 |
Read the first row again. More than 17,000 requests hunting for secrets, and every one arrived announcing itself as a crawler you would normally welcome. Not one credential scan came from a verified operator. Nobody at Google, OpenAI, or ByteDance is looking for your environment file. Attackers wearing their names are.
The unverifiable column matters as much as the impostor column. A request in that bucket claimed a trusted identity that could not be confirmed, and more than half of all attack traffic on the network sits there. That is why an unconfirmed claim is never counted as real. Treating it as friendly is how the credential scanner gets waved in.
Whose Names the Fakes Steal
Ranked by failed-verification visits across the fleet, as of September 2026:
| Claimed identity | Fake visits | Unique sources |
|---|---|---|
| Bytespider | 11,298 | 444 |
| Amazonbot | 6,966 | 170 |
| Googlebot | 2,629 | 628 |
| Applebot-Extended | 747 | 31 |
| BingBot | 615 | 97 |
| Applebot | 440 | 83 |
| Meta-ExternalAgent | 341 | 63 |
| TikTokSpider | 283 | 155 |
| DuckDuckBot | 180 | 63 |
| FacebookBot | 113 | 3 |
Two things stand out. Bytespider and Amazonbot together account for more than three quarters of all forged visits, so if you have ever looked at a log and thought Bytespider was hammering you, a good share of that was never ByteDance. And Googlebot is forged from more distinct sources than any other name, 628 of them, because Googlebot is the name most likely to be trusted on sight.
Every entry in that table is a request that claimed a real crawler's identity and failed to prove it. The real Googlebot is welcome. The 2,629 requests pretending to be Googlebot are the ones you never want near your login page.
Zero Percent of Our Bytespider Was Real
The finding that started this work still holds. In the first full verification pass across our own network, every source claiming to be Bytespider failed. Fourteen distinct sources, zero verified, 100 percent impostor. Not one came from ByteDance.
The fleet-wide number above shows it was not a fluke. Bytespider is the most-forged crawler name on the internet for a simple reason: everyone already blocks it, so nobody looks closely at it.
One Source Wore Seven Names
A single address, 45.45.237.69, appeared in the logs claiming to be seven different AI crawlers: BingBot, ChatGPT-User, ClaudeBot, GPTBot, OAI-SearchBot, PerplexityBot, and Google-CloudVertexBot. One source, seven of the most trusted names in AI, none of them real.
No category chart catches that. It only shows up when you stop counting names and start asking each request to prove who it is. That one row is the entire argument for verification.
What an Attack Looks Like When It Hides Behind a Crawler's Name
On the morning of August 27, twenty sources hit one client site in a seven-minute burst. About 170 requests, 55 distinct URLs. Every source rotated between the Googlebot and BingBot user-agents from one request to the next. Every one went after environment files, with encoded slashes, null bytes, and doubled path separators to slip past filters that match on the obvious spelling.
Every one of the twenty was flagged as an impostor the moment it arrived. Nothing sensitive was served. But if you were skimming raw logs and trusting the user-agent, you would have recorded twenty friendly visits from Google and Bing and moved on.
That is the whole point. Claimed identity, verdict, and behavior are three separate facts, and the middle one is the only thing that keeps the first from lying to you about the third.
Why This Belongs in Your AEO Data, Not Just Your Security Review
Verification is a scalpel, not a panic button. Across the network the verdicts come back roughly 84 percent verified, 6 percent impostor, and 9 percent unverifiable. Most crawler traffic is exactly what it says it is, and the real operators are reading your pages, not probing them.
But if 1 in 14 of the entries in your "AI crawler" report is a forgery, your picture of who is reading you is wrong before you start optimizing. The fakes inflate your Bytespider count, pollute your Googlebot count, and put a credential scanner in the same row as OpenAI. Verify every claimed crawler first, set the impostors aside, and only then do you have numbers worth acting on. For the wider problem of sorting real agents from noise, see how to classify AI crawler traffic without guessing. If you decide some of it should not reach you at all, how to block AI crawlers covers the trade-offs.
Read the Source, Not the Chatbot
Most AI-visibility tools never see any of this, because they work by asking a model questions and guessing whether your brand showed up. They never touch your server, so they cannot see a fake Bytespider hunting your .env, or one source wearing seven names. Tools that read logs but stop at the user-agent are only slightly better: they count every forgery as a real visit because the forgery carries a real name.
For WordPress sites, citAEOtion tracks real AI crawler visits at the server, verifies every claimed identity, and reports what the verified crawlers read and what the impostors reached for, per crawler, per page, and per site. It is the same data that told us why Copilot citations dropped in late July. Pricing starts at $34.99 per month, or book a 20-minute walkthrough to see the verdicts on your own traffic. You cannot verify what you refuse to check, and you cannot check what a chatbot never saw.
Frequently Asked Questions
What are fake AI crawlers?
Requests that carry the user-agent of a real crawler such as Googlebot, Bytespider, or GPTBot but do not come from the operator behind that name. The user-agent is set by the sender, so the name alone proves nothing. Across the sites citAEOtion monitors, 23,614 requests have failed verification as of September 2026.
What do fake AI crawlers actually do?
A third of them probe for credentials: environment files, cloud metadata, configuration and deploy secrets. The rest scrape content or scan for weaknesses. None of them send you visitors or citations.
Does a fake Googlebot mean Google attacked my site?
No. It means a request claimed to be Googlebot and could not prove it. The real operator and the impersonator are different senders. citAEOtion records the claimed identity, the verdict, and the behavior as separate facts and never attributes a forgery's actions to the operator it imitated.
How do I know which crawler requests are real?
Each claimed identity has to be verified rather than trusted, and the result is a verdict on every request: verified, impostor, or unverifiable. citAEOtion does this automatically for every crawler that claims to visit you, so the crawler traffic in your reports is the real thing and the rest is set aside with a record of what it tried.
Does blocking fake AI crawlers hurt AI visibility?
No. Impostors send no traffic and no citations, so removing them only removes attackers and noise. Verified crawlers from real operators are left untouched, which is the reason to separate the two before you block anything.