AI Research

Pew study finds signs of AI authorship across a large share of newer web pages

A Pew Research Center analysis of nearly 500,000 English-language web pages finds that AI-authored text is increasingly visible online, especially on pages published after ChatGPT’s launch.

Published Updated
Pew Research CenterAI WritingWeb Content

A new Pew Research Center study offers one of the clearest measurements yet of how much AI-written text has entered the public web. Published on August 20, the analysis examined nearly 500,000 English-language pages collected from Common Crawl over a five-year period and used a detection approach called Open Pangram to estimate whether pages showed significant signs of AI authorship. Pew reported that, in a random sample of 10,000 pages from July 2026, about 10 percent showed meaningful evidence of AI-generated writing.

The share is much higher for newer material. Pew found that among pages with a publication date after the launch of ChatGPT, more than one-third showed indications of AI authorship. That does not mean every sentence on those pages was written by a model, and it does not prove that human editors were absent. It does suggest that generative AI has become a common part of web publishing in only a few years, especially on pages created after large language models became widely accessible.

The study also found large differences by domain type. Pages on commercial dot-com domains showed AI signals more often than nonprofit, educational or government domains. Pew reported that roughly one in ten dot-com pages in the July 2026 sample showed signs of AI authorship, compared with a smaller share of dot-org pages and very low levels on dot-edu and dot-gov sites. That pattern fits the economics of web publishing, where commercial sites have stronger incentives to produce large volumes of search-friendly text, product pages, marketing copy and low-cost informational content.

Pew’s work is careful about uncertainty. AI detection is not perfect, and the center does not claim to identify every use of AI or every mixed human-machine editing process. Instead, the study uses repeated patterns in language to estimate large-scale change. It also tracks linguistic shifts that became more common after ChatGPT’s release, including changes in punctuation, style and certain words or sentence structures. Those clues are useful at web scale even if they should not be treated as proof in a single document.

The findings matter because the web is both a publishing platform and training data for future AI systems. If more pages are written with AI, search engines, researchers and model developers have to think carefully about feedback loops. AI-generated material can be helpful when it explains a topic clearly, localizes information or helps small organizations publish. It can also flood search results with repetitive, thin or synthetic text that makes it harder to find original reporting, expert analysis and firsthand documentation.

For publishers, the study sharpens an uncomfortable question: what counts as quality when AI assistance becomes ordinary? Using a model to edit, translate or structure material is not the same as running automated content farms. The problem is not AI use by itself, but the absence of human judgment, sourcing and accountability. Readers rarely object to tools that make good work clearer. They do object when pages appear designed mainly to capture search traffic without adding knowledge.

The larger implication is that AI literacy now includes source literacy. Users need to ask not only whether a page is accurate, but how it was made, whether it cites evidence and whether it reflects real expertise. Pew’s study does not declare the web broken. It shows that the texture of the web has changed quickly. Search, journalism, education and AI training pipelines will all need better ways to distinguish useful assisted writing from synthetic volume that merely looks like information.