824 IPs Pretending to Be GPTBot and ClaudeBot Are Hunting Your .env Files: Inside GreyNoise's Fake AI Crawler Campaign
GreyNoise says scanners forged the identities of 13 AI crawlers from 824 addresses to request .env files, cloud keys and password stores — and not one address matched the published ranges of OpenAI, Anthropic, Google, Perplexity or Amazon.
Here is a sentence that should unsettle anyone who runs a website in 2026: your logs almost certainly contain traffic from GPTBot, ClaudeBot and Google’s crawlers, and for a four-week window this summer, some of that traffic was lying about who it was.
On August 28, 2026, threat-intelligence firm GreyNoise published an investigation that reads like a phishing campaign aimed at infrastructure rather than inboxes. Between July 28 and August 23, a cluster of automated scanners forged the user agent strings of 13 AI crawlers belonging to eight companies — including OpenAI, Anthropic, DeepSeek, Google and Perplexity — and used those stolen identities to request the files no legitimate crawler would ever ask for: .env configuration files, cloud access keys, private keys and password stores.
The scale is precise. Six of the forged crawler names, belonging to four AI companies, arrived from the same 824 IP addresses in almost identical volume, all riding a single HTTP client fingerprint. That fingerprint, GreyNoise noted, had carried more than 1,500 different user agent strings over the preceding 90 days — most of them masquerading as ordinary browsers. The largest single day of activity was August 23, suggesting the campaign was still accelerating when the observation window closed.
Why the disguise works
The mechanics of the trick are almost embarrassingly simple. Every program that visits a website announces itself in one line of the HTTP request — the user agent. Chrome says it is Chrome; Googlebot says it is Googlebot; Anthropic’s crawler says it is ClaudeBot. But as the GreyNoise researchers put it, nothing in the request itself proves any of it is true. The user agent is a client-supplied header, and any script can set it to any string in milliseconds.
The reason this particular lie is dangerous now is that AI companies have spent two years persuading site owners to whitelist their crawlers. Publishers who want their content in ChatGPT answers add GPTBot to their allowlists. WAF rules, bot-management policies and rate-limit exemptions are increasingly keyed to crawler names. The fake-crawler campaign weaponizes that trust: a control that checks the name but not the source address “can be bypassed by forging it,” as the researchers wrote.
And the forgery is careful. The impostors’ ClaudeBot string matches Anthropic’s published user agent character for character. No rule keyed on the user agent can tell the two apart, because the label itself is identical.
How GreyNoise knows they’re fake
Three independent lines of evidence give the game away.
First, the behavior is wrong. A real crawler reads /robots.txt before almost anything else — it is the file where a site states its rules, and Anthropic’s genuine crawler requested it more than any other path, accounting for 12% of its traffic over the same window. The six forged names never requested /robots.txt once. What they asked for instead were credential paths: /.env, /app/.env, /api/.env, /backend/.env, /.env.local, /.env.production, /.env.old, /.env.bak, /.aws/credentials, even /.env.swp — a vim swap file. Across all traffic on this fingerprint, requests for secrets “ran into the millions.”
Second, the addresses don’t check out. All four AI companies (and Amazon) publish the IP ranges their crawlers use — Anthropic at claude.com/crawling/bots.json, OpenAI at openai.com/gptbot.json, and so on. GreyNoise fetched every one of those lists and tested every one of the 824 addresses against them. Not a single address matched. Meanwhile, over the same period, thousands of sessions carrying the ClaudeBot name did arrive from Anthropic’s actual published ranges — the real crawler and the impostor were running concurrently, distinguishable only by source address.
Third, one name was impossible. The scanners also sent 263,849 sessions carrying “Google-Extended” — a token publishers write in robots.txt to opt out of AI training. Google explicitly documents that it has no separate HTTP request user agent string. No Google crawler sends it. Every one of those 263,849 sessions was forged by definition.
The actors also forged two of Amazon’s crawler names, in even greater volume than the six matched AI names — under user agent strings that Amazon does not document at all.
Why blocking is hard
The natural response — block the IPs — fails here by design. The 824 addresses are spread across 795 separate /24 networks, so there is no single network or ASN to block. GreyNoise published the full address list, along with every observed credential path and the (half-redacted) impostor JA4H fingerprint, ge11nn05enus_f3bb7a..., recommending it for investigation rather than blocking.
GreyNoise was also careful about what it does not claim: nothing in the data says any file was actually returned, or that any named organization was affected, and the company is not attributing the activity to anyone.
What to actually do
GreyNoise’s recommendations split neatly by audience, and they double as a checklist for the AI era’s newest trust problem:
- Security operations: never treat a user agent string as identity. Check the connecting address against the published list for the name it claims — each crawler has its own list, so check GPTBot traffic against
gptbot.json, not a merged blob. Alert on any request for/.env,/.aws/credentialsor/.git/config; no crawler has any reason to ask for these. And judge robots.txt behavior across days, not single visits, since real crawlers cache the file. - Security leadership: find every place a user agent string grants access or waives a control, and put a real check behind it. Give each vendor’s address list an owner and a refetch schedule — a stale list turns the real crawler into an alert.
- Web and platform admins: keep
.env,.gitand cloud credential files out of the web root entirely. Rotate any cloud key that was ever reachable from a web path, and assume anything readable was read. The researchers also flagged patching Vite to the fixed releases (6.2.3, 6.1.2, 6.0.12, 5.4.15, 4.5.10) to close the arbitrary-file-disclosure hole that makes these probes land harder.
The bigger picture
This campaign lands on an internet mid-way through a messy renegotiation of who gets to crawl whom. AI companies want access; publishers want payment or exclusion; robots.txt — a 1994 convention built on gentleman’s agreement — is suddenly load-bearing infrastructure for the AI economy. The GreyNoise findings show what happens when that handshake becomes a target: the same allowlists built to welcome AI crawlers become an attack surface, and the verification lists the AI companies publish become the only thing standing between a site owner and a credential-harvesting scan wearing a friendly name.
There is a bitter irony in the tell. The impostors were caught not by fancy fingerprinting but by the oldest norm on the web — good bots read robots.txt first. The fakes skipped the courtesy and went straight for the keys. In an era when AI agents are increasingly allowed through the front door on their stated name, the lesson generalizes far beyond crawlers: identity you can assert in a header is not identity. Verify the address, or you are trusting a string.
Sources
- [1] https://www.greynoise.io/blog/threat-actors-posing-as-ai-crawlers
- [2] https://www.helpnetsecurity.com/2026/08/31/ai-crawlers-scan-exposed-credentials/
- [3] https://ppc.land/scanners-posing-as-claudebot-hunt-credential-files-from-824-addresses/
- [4] https://ground.news/article/threat-actors-are-posing-as-ai-crawlers-to-hunt-for-exposed-credentials