← All posts / Research

One in Ten Web Pages Is Now AI-Written, and Among New Pages It's One in Three: Inside Pew's Landmark Web Study

Pew Research Center analyzed 490,000 English-language web pages and found 10% show significant signs of AI authorship as of July 2026 — a share that climbs past 35% among pages published since ChatGPT's launch, with measurable shifts in punctuation and vocabulary across the entire web.

One in Ten Web Pages Is Now AI-Written, and Among New Pages It's One in Three: Inside Pew's Landmark Web Study

How much of the internet is actually written by machines? For years, that question lived in the realm of internet folklore — the “dead internet theory,” a conspiracy-adjacent claim that most online content is bot-generated, was easy to dismiss because nobody had rigorous numbers. On August 20, 2026, Pew Research Center’s Data Labs team published the most credible answer yet: in a random sample of 10,000 English-language web pages collected in July 2026, roughly one in ten shows significant signs of being written or substantially edited by AI. And among pages published since ChatGPT’s November 2022 release, the share crosses one in three.

The study, led by Pew senior data scientist Samuel Bestvater with Aaron Smith, Carson TerBush, Chris Baronavski, and Janakee Chavda, is the largest systematic audit of AI authorship on the public web to date — and its fine-grained findings tell a more unsettling story than the topline number alone.

How the study worked

Pew drew its corpus from Common Crawl, the open web-scraping archive that underpins much of the research world’s picture of the internet. The team collected almost half a million English-language web pages — about 490,000 — spanning January 2021 through July 2026. That starting point matters: it establishes a pre-ChatGPT baseline roughly two years before the chatbot era began, letting the researchers watch the AI signature emerge in the data rather than infer it.

To classify pages, Pew ran the text through Open Pangram, an open-weight AI detection model built by detection startup Pangram. The model doesn’t hunt for a single smoking gun. It looks for statistical patterns in word choice, phrasing, and sentence structure that distinguish machine-generated prose from human writing — linguistic “tells” that are individually weak but collectively diagnostic at scale.

Pew is candid about the method’s limits: AI detectors sometimes misclassify individual documents in both directions. But as the report notes, when you aggregate very large collections of texts, systematic differences between human and machine writing rise above the noise. That aggregation-level honesty is what makes the study more defensible than the detector-driven controversies that have plagued academic integrity enforcement.

The numbers: a web visibly transforming

Three layers of findings stand out.

The snapshot. Of 10,000 randomly sampled pages from July 2026, 10% showed significant signs of AI authorship. That may sound modest — until you remember the denominator includes every aging page from the pre-AI web that is still indexed.

The trend. Filter to only pages published after ChatGPT’s launch, and AI authorship signals appear in over 35% of new pages in the July 2026 snapshot. The trajectory has climbed steadily since late 2022 with no plateau in sight — each successive sample shows a higher rate than the last, tracking the succession of chatbots from ChatGPT through Claude and Gemini.

The domain split. AI writing is not evenly distributed. Around one in ten .com pages shows AI authorship — roughly double the rate on .org domains (4.6%) and ten times the rate on .edu and .gov domains, which both sit near 1%. The chart data shows the divergence clearly: .com pages jumped from 1.09% AI-flagged in January 2021 to 9.35% by January 2026, while .edu and .gov barely moved from sub-1% baselines. Commercial incentives, it turns out, are the strongest predictor of synthetic content.

The linguistic fingerprints

The study’s most quietly remarkable section quantifies how AI has literally changed the way the web writes. Comparing today’s internet to a 2023 snapshot, Pew found:

  • Em dashes appear about twice as frequently — the punctuation mark of AI’s journalistic training data, now colonizing human-adjacent prose.
  • Oxford commas saw a 63% increase in usage per 10,000 words, rising from 34.0 in early 2023 to 55.5 by January 2026.
  • AI-favored vocabulary has more than doubled. Words like “delve,” “interplay,” “testament,” “underscore,” “pivotal,” and “tapestry” — the full tracked list includes 27 terms from “additionally” to “vibrant” — climbed from 11.9 per 10,000 words to 26.0.
  • Negative parallelism (“it’s not just X, it’s Y”) nearly tripled, though it remains rare in absolute terms (0.87 to 2.36 per 10,000 words).

The data table behind the vocabulary chart is worth pausing on: em dash density went from 5.79 per 10,000 words in January 2023 to 11.19 by January 2026. These aren’t stylistic abstractions — they are measurable population-level shifts in the texture of English prose online, driven by machine preferences absorbed from training data and now reflected back at scale.

Pew is careful to note that any individual page with an em dash or a “delve” is obviously not proof of AI authorship — humans use these too. The signal only becomes meaningful in aggregate, which is exactly why the researchers frame it as a statistical fingerprint rather than a per-document verdict.

Why it matters

For training data. The AI industry’s foundational assumption — that the web offers an ever-growing supply of fresh human text — is weakening in real time. As post-2022 pages increasingly contain machine output, model-collapse and data-contamination concerns shift from theoretical to measurable. The .com rate of ~10% (and rising) is precisely where crawlers harvest most aggressively.

For trust and provenance. The domain split offers a rare piece of good news: institutional domains (.edu, .gov) have remained largely resistant, suggesting that editorial review pipelines and accountability structures still gate synthetic content effectively. That contrast — commercial web flooding, institutional web holding — may become the defining provenance divide of the next decade.

For the dead internet debate. CNET and other outlets framed the study against the dead internet theory, and the verdict is nuanced: the theory’s strong form (most content is bots) remains false, but its directional instinct is vindicated — a third of new commercial web pages carrying AI fingerprints is no longer fringe paranoia. With half of U.S. adults now using AI chatbots and 24% using them daily, per Pew’s own survey data, the supply side and demand side of synthetic text are compounding each other.

For detection skepticism. Pew’s decision to publish alongside a full methodology page, lean on aggregate-level claims, and acknowledge detector error rates models the epistemic humility the field needs. It also lands amid an active debate — The Verge’s recent coverage documented how detector companies like Pangram claim false-positive rates as low as 1 in 10,000 while independent analyses sometimes find 2% — a gap this study doesn’t resolve but partially sidesteps by never claiming per-page certainty.

The bottom line

Pew has turned a meme into a measurement. Ten percent of the sampled web — and over a third of the post-ChatGPT web — now carries machine fingerprints, and the linguistic substrate of English online is bending measurably toward model preferences. The open questions are the important ones: whether the curve keeps climbing once detection arms races play out, whether regulatory provenance regimes (watermarks, C2PA-style provenance) gain traction, and whether the .edu/.gov holdout zones survive as AI-assisted writing normalizes. What is no longer debatable is the direction. The web that trained the models is being rewritten by them.