← All posts / Industry

ChatGPT Stopped Reading the Web's Middle Layer: An 85% Retrieval Collapse, Two Datable Cliffs, and What Publishers Can Still Fix

Fifty-seven days of server logs show ChatGPT retrievals to one AI-tools site falling 85% — from 7,507 on August 5 to 1,118 on September 18 — while OpenAI's crawl of the wider web roughly tripled. The curve has two datable cliffs: the August 8 shift to domain-scoped retrieval and Cloudflare's September 15 AI-crawler defaults.

ChatGPT Stopped Reading the Web's Middle Layer: An 85% Retrieval Collapse, Two Datable Cliffs, and What Publishers Can Still Fix

For two years, publishers tracked AI traffic with a single anxious question: how much is ChatGPT sending us? The question this month is stranger and more urgent: why has ChatGPT stopped reading us — while crawling the rest of the web harder than ever?

On September 20, AIToolsRecap, an independent AI-tools publication, published a log analysis that is quietly becoming one of the most-cited data points in the AI-visibility debate. Across 57 days of its own server logs, ChatGPT retrievals fell from a peak of 7,507 on August 5 to 1,118 on September 18 — an 85 percent collapse. Over the same window, OpenAI’s crawl of the wider web roughly tripled. The web is being crawled more; this site, and by the evidence an entire category of sites, is being retrieved less.

What it is not

The analysis is careful to rule out the obvious suspects before assigning blame.

It is not OpenAI crawling less. Across the wider web, OAI-SearchBot activity is up roughly 3.5x and GPTBot roughly 2.9x since the GPT-5 launch, corroborated by third-party log analyses covering billions of log lines.

It is not robots.txt. The site’s file is User-agent: * with Allow: /, a sitemap, and no disallow lines. Nothing was blocked there.

And it is not a smooth decay — which is the most useful diagnostic detail in the data. Gradual decline would suggest content going stale. Step functions suggest that somebody, somewhere, changed a setting. The curve has two cliffs with flat stretches between them.

Cliff one: August 8, domain-scoped retrieval

On August 8, ChatGPT Search’s use of the site: operator inside its fan-out queries jumped from about 0.37 percent to 16.8 percent in a single day. Instead of searching the open web and choosing among results, the system began scoping queries to specific domains — asking what a particular site says about a topic — and favoring official brand sites in what it cited.

The corroborating number is Reddit. Between August 8 and 17, Reddit’s share of ChatGPT citations fell from a 3.83 percent average to 0.52 percent — an 86 percent relative drop. Reddit and a tools directory are the same kind of source: third parties writing about other people’s products. When retrieval scopes to the vendor’s own domain, that is exactly the category that gets skipped. The publication’s 85 percent and Reddit’s 86 percent are, in the analysis’s blunt phrasing, “the same event.”

A related contraction happened just before it: between late July and early August, fan-out queries in paid Thinking mode fell from 3.56 to 1.90 per conversation, and URLs touched fell from 47.2 to 33.9. Fewer pages per answer — then a preference shift in which pages.

Cliff two: September 15, Cloudflare’s defaults

On September 15, Cloudflare changed its AI-crawler defaults. Training and Agent crawlers are now blocked by default; Search bots remain allowed. The change applied to new sites and to existing free-plan customers still on default settings who had not set a preference beforehand.

The logs show the effect immediately: 1,907 retrievals on September 16, 1,382 on September 17, 1,118 on September 18 — a 41 percent fall in two days. The analysis is blunt about the remediation: if you are on Cloudflare and never opened the AI settings, “this is the one you can undo this afternoon.”

The third factor nobody can fix

Live page fetches only happen in Thinking mode. Free instant mode does not open pages at all — which the analysis characterizes as an economics decision at OpenAI, not a caching artifact. Retrieval now runs mostly off a discovery index and a shared reading cache, and the live fetch is reserved for the mode where users pay and wait. Industry-wide, live fetches reportedly fell about 28 percent between December 2025 and March 2026 — before either cliff.

Three bots, three different jobs

The most actionable section concerns a mistake most analytics dashboards make: lumping “ChatGPT” into one column. It is three different user agents doing three different jobs.

GPTBot collects training data; Cloudflare now blocks it by default, and commercially it costs a publisher little. OAI-SearchBot builds the ChatGPT search index; it remains allowed by Cloudflare’s new default, so a fall there is a standing problem, not a blocking one. ChatGPT-User performs the live fetch during thinking-mode sessions and is largely outside a publisher’s control.

The recommended procedure is to split the column before changing anything — and to test with a real user-agent string and read the status code directly, because Cloudflare’s own dashboard bundles blocked, rate-limited, and not-found responses into a single number that will mislead you.

The Bing asymmetry

One more number deserves attention. Over the same 57 days, the site’s logs recorded 82,976 Bingbot hits against 12,927 Googlebot hits. Bing crawls the site 6.4 times as often as Google. That asymmetry goes a long way toward explaining why the same pages that Copilot cites constantly are nearly invisible in Google search — a page cannot rank in an index that rarely looks at it.

What the publisher is doing about it

The response is a playbook other sites can copy: split the user agents, check the Cloudflare toggle — the cheapest possible fix if it applies — accept the August loss, because domain-scoped retrieval structurally disadvantages anyone writing about products they do not own, and write more of what a vendor cannot. An official site can publish its own pricing; it cannot publish a tested head-to-head against a competitor, and it cannot publish a first-party log analysis like this one.

The honest caveat

The analysis is candid about its limits: one site’s logs, skewed toward AI tooling coverage — precisely the category most exposed to a shift toward official sources. “A single domain is an anecdote, not a study.” Treat the dates as well-evidenced and the magnitudes as specific to that site. But the two cliffs should show up in any publisher’s data across those dates; if they do not, the explanation is wrong.

Why this matters beyond one site

The AI-visibility story of 2026 has mostly been told in citation counts and referral-traffic surveys. This analysis adds something those lack: dated, mechanical causes with named thresholds. If domain-scoped retrieval and crawler-default changes can erase 85 percent of ChatGPT’s reading of a category of sites in six weeks, then AI visibility is not a ranking problem — it is a policy problem, decided in settings screens at OpenAI and Cloudflare that most publishers have never opened.

For anyone who publishes on the web, the checklist is short. Split your bot logs. Check your Cloudflare AI settings. And recognize that the content that still gets retrieved is the content an official source cannot publish about itself — which may be the last durable moat in the AI search era.