← All posts / Industry

Tomorrow the Web's Default Setting Flips: Cloudflare's Content Independence Day Arrives

On September 15, Cloudflare blocks Training and Agent crawlers by default on ad-bearing pages for all new domains — the largest single shift in crawler economics ever attempted, and a wake-up call for mixed-use bots like Googlebot.

Tomorrow the Web's Default Setting Flips: Cloudflare's Content Independence Day Arrives

For thirty years, the web ran on an unwritten contract: crawlers take your pages, and in exchange you get referral traffic. That deal died when AI answer engines began absorbing content and sending back nothing. Tomorrow — September 15, 2026 — Cloudflare formalizes the funeral. For every new domain that onboards to its network, Training and Agent crawlers will be blocked by default on pages that display ads, while Search crawlers stay allowed. It is the largest single intervention in the economics of automated traffic ever attempted, and it lands on roughly a fifth of the web at once.

What exactly changes on September 15

Cloudflare first announced the shift back on July 1, 2026 — its second annual “Content Independence Day” — and gave the industry a 75-day runway to prepare. The mechanics, confirmed in the company’s blog post and developer changelog, are straightforward:

  • New domains onboarding to Cloudflare get the new defaults automatically: Training blocked, Agent blocked, Search allowed — specifically on pages that serve ads.
  • Existing customers keep their current configuration, but Cloudflare has been notifying them as the date approached; anyone who wants different rules must set them explicitly in Security settings.
  • Multi-purpose crawlers — bots that combine Search with Training under one user-agent — will now be treated according to all of their behaviors, enforced by the most restrictive applicable rule.

That last bullet is the sleeper clause. Because Googlebot, Applebot, and BingBot all blend search indexing with AI training under shared infrastructure, any site owner who blocks Training will also block those mixed-use crawlers entirely. Cloudflare’s position is deliberate: bot operators were told to separate their crawlers by purpose, and the ones that didn’t are now paying the price of opacity.

The taxonomy behind the flip

The change isn’t just a block button. Cloudflare spent the past year rebuilding its classification system around a pragmatic taxonomy of behavior rather than the fuzzy label “AI”:

  • Search — proactively builds a database of your site to answer queries later. Site owners should expect referral traffic or equitable compensation.
  • Agent — acts in real time on a person’s behalf to complete a job, from ChatGPT-User fetches to Gemini or Claude driving a full browser.
  • Training — permanently absorbs your content into a model’s weights.

Under the old model, a small site faced a Faustian bargain: allow everything and get trained on without compensation, or block everything and vanish from discovery. The new three-way split dissolves that dilemma — and it’s available to Free-tier customers, not just enterprises.

Verified status now has teeth

Alongside the defaults, Cloudflare redefined what “Verified bot” means. Previously, Verified implied default-allowed. Now it means allowable within its category — a bot is permitted only if the site owner has opened the door for that category of behavior. Verification also carries a enforcement stick: any bot caught abusing content-use signals, or reproducing content in full, loses Verified status and with it access across the more than 20% of web domains behind Cloudflare.

A new BotBase database gives Enterprise customers a searchable dashboard of every known bot, its classification, and its content-use policy, with detection IDs copyable straight into Security rules. And a new robots.txt signal — use=immediate, use=reference, or use=full — extends Content Signals so site owners can express how much of their content a bot may keep and reshare, from “interact but store nothing” to “summarize and reproduce.”

The agentic internet gets an identity layer

The most forward-looking piece addresses transitive trust: when an agent running on a developer platform calls your API through three layers of intermediaries, who do you actually trust? Cloudflare proposes using the existing RFC 7239 Forwarded header to carry operator identity end-to-end:

Forwarded: for="openai";use="reference"

A site owner who allows OpenAI’s agents can keep that preference intact whether the agent arrives directly or through intermediaries — and an operator that abuses the trust loses it network-wide. Trust becomes portable, and revocable.

Why this matters

The scale is the story. Cloudflare proxies over 20% of all web domains, and its Radar data already showed bots crossing the majority line of HTML traffic earlier this year — machines now outnumber humans on the web. When the default flips tomorrow, the burden of permission inverts: AI companies can no longer assume access to ad-funded content; they must negotiate it, separate their crawlers, or get blocked as a baseline state.

For publishers, it’s the first real structural counterweight to answer-engine economics — a default that favors compensation over extraction, paired with the Pay Per Crawl marketplace launched last year as the monetization rail. For AI companies, it’s a compliance deadline that arrived with teeth: mixed-use crawling is no longer a loophole but a liability. And for the web as a whole, it’s a live experiment in whether default settings — not lawsuits or regulation — can rewrite a broken bargain between content makers and content takers.

The 75-day grace period ends at midnight. The web’s most consequential A/B test on AI access begins then.