Analyzing nearly half a million pages, Pew Research found roughly 35% of pages published after ChatGPT's late-2022 release carry AI-authorship signals — but only about 1 in 100 on .edu and .gov domains.
2026-08-20 · by Moe Ameen
On August 20, 2026, the Pew Research Center published a study estimating how much of the web now shows signs of AI authorship. Using the Common Crawl archive, researchers analyzed nearly 490,000 English-language web pages spanning January 2021 through July 2026 — a window that starts well before and runs long after ChatGPT's public release in November 2022. To flag AI writing, they ran the pages through Open Pangram, a detection model that looks for the statistical language patterns — word choices, phrasings, and stylistic quirks — that AI systems use more often than human writers do.
The headline figure: over one-third — about 35% — of pages published after ChatGPT's release show signs of AI authorship. Across a random sample of 10,000 pages collected in July 2026, roughly one in ten showed those signs overall, and the rate varied sharply by domain type. About 10% of .com pages carried AI signals, roughly double the 4.6% share on .org pages and about ten times the rate on .edu and .gov pages, which both sat near 1%. Pew also tracked the fingerprints themselves: since 2023 the frequency of em dashes has roughly doubled, Oxford commas are up about 63%, and words like "delve" and "pivotal" alongside "it's not X, it's Y" negative-parallelism phrasing have become common tells.
Pew is careful about what the number does and does not mean. This is not a claim that a third of the entire internet was written start-to-finish by a machine — it is a measure of pages showing signs that AI played a role in writing or heavily editing them. And detection is imperfect: like every AI detector, Open Pangram sometimes misclassifies human writing as AI and vice versa. The researchers argue that at the scale of hundreds of thousands of documents the aggregate pattern is reliable even though any single page's label may be wrong, so treat the figure as directionally accurate rather than a precise per-page verdict.
Read the Pew number as a map of what everyone else's content now looks like: a third of new pages statistically averaged, carrying the same handful of tells a detector is trained to count. The way out is not a "humanizer" that games the checker — those get caught too — it is content that is genuinely in your voice and, increasingly, not plain text at all. That is the specific gap [Kompozy](/) is built to close. Its [Persona Brief](/glossary/persona-brief) encodes your actual voice and runs a banned-word filter that strips the exact clichés this study catalogs — the reflexive "delve," "pivotal," and "it's not X, it's Y" negative parallelism — so what ships reads like you, not like the web's new average.
The bigger lever is format. Pew measured written pages, and the surest way to escape the text-sameness trap is to stop competing only on text. From one idea, Kompozy generates captioned [Persona Shorts](/glossary/persona-shorts) and avatar video in your own likeness, brand-exact [Carousels](/glossary/hyperframes), Photo Posts and Quote Graphics, plus a Blog Article and an Email Newsletter — native assets a text detector never even looks at, each carrying a face and a voice a model cannot average away. Every piece passes a per-post human review gate before [Autopilot](/glossary/autopilot) schedules and fans it out across the eight social platforms plus blog and email. If you want the deeper context on how these detectors work and why they misfire, our guide to [AI content detection](/guides/ai-content-detection) and the note on [Pangram's detection funding](/news/pangram-9m-ai-detection-funding) — the tech behind this very study — are the next reads; for the differentiation playbook, see [AI content saturation on social media](/guides/ai-content-saturation-social-media).
No. It found that over one-third — about 35% — of pages published after ChatGPT's release show signs of AI authorship, meaning AI likely helped write or heavily edit them. That is different from a page being generated start-to-finish by a machine. Across a July 2026 sample of all pages, roughly one in ten showed AI signals overall, and Pew stresses the number is directionally reliable rather than a precise per-page verdict.
Pew ran nearly 490,000 English-language web pages from the Common Crawl archive through Open Pangram, a detection model that flags the statistical language patterns AI systems use more than human writers — specific word choices, phrasings, and stylistic quirks. Like all AI detectors it misclassifies some pages in both directions, so Pew treats the aggregate across hundreds of thousands of documents as reliable while acknowledging any single label may be wrong.
In Pew's July 2026 sample, about 10% of .com pages showed signs of AI authorship — roughly double the 4.6% on .org pages and about ten times the rate on .edu and .gov pages, which both sat near 1%. The pattern suggests AI writing is most common on commercial sites and rarest where trust and institutional authority matter most.
Compete on voice and format, not on more text. Write in a distinct, recognizable voice that avoids the tells a detector counts — reflexive em dashes, "delve," "pivotal," and "it's not X, it's Y" phrasing — and lean into formats a text detector never sees, like captioned video, avatar shorts, and brand carousels. A tool like Kompozy governs voice with a Persona Brief and banned-word filter and turns one idea into native video, image, and long-form assets, with a human review step before anything publishes.