Two credible 2026 studies put wildly different numbers on the same question. Pew Research analyzed nearly half a million web pages and found that about one in ten shows signs of AI authorship — rising to over a third among pages published since ChatGPT launched. Graphite, an SEO firm, sampled newly published articles and found the AI share sitting near half. Ten percent or fifty percent is a big gap, and the instinct is to decide which study is wrong. Neither is. They measured different populations, with different detectors, against different thresholds, and once you line those up the two numbers describe the same reality from different angles. This guide is a statistics-literacy piece for anyone who publishes: what the Pew figure actually counted and how it was measured, why the domain breakdown (commercial .com pages carry roughly ten times the AI rate of .edu and .gov) is the most useful part of it, why Graphite's number is higher, the four questions that reconcile any 'X% of the web is AI' headline, and the practical conclusion — that the web average is not your problem to solve, being the exception these studies don't count is.
If you have read anything about the state of the web in 2026, you have seen two figures that seem to contradict each other. One says roughly a tenth of web pages show signs of being written by AI. The other says about half of new articles are machine-made. Ten percent and fifty percent are not close, and the natural reaction is to assume one study is sloppy. That reaction is wrong, and correcting it is the single most useful thing you can learn about AI-content statistics. Both numbers come from credible work; they simply answer different questions. Once you know which question each one is answering, they stop competing and start reinforcing each other.
The lower figure is Pew Research's. The higher figure is Graphite's. The gap between them is not error — it is the difference between counting all the pages on the web and counting only the fresh articles being published right now. This guide walks through exactly what each study measured, why the domain breakdown inside the Pew data is the part worth memorizing, and the short checklist that lets you read any future 'X% of the internet is AI' headline without being fooled by it. For the deeper single-study treatment of the newly-published-articles view, the companion piece on how much of the web is AI-written covers the Graphite research and the distribution finding in full; this guide is the one that reconciles it with the broader page-level numbers.
On August 20, 2026, the Pew Research Center published an analysis of how much of the web now shows signs of AI authorship. The method matters, because it defines the number. Researchers took nearly 490,000 English-language web pages from Common Crawl — a large, public archive of the web — spanning January 2021 through July 2026, a window that starts well before ChatGPT's public release and runs long after it. They ran each page through Open Pangram, a detection model that flags the statistical language patterns AI systems use more often than human writers: certain word choices, phrasings, and stylistic quirks.
Two headline figures came out of it. In a random sample of 10,000 pages collected in July 2026, about 10% of all pages showed signs of AI authorship. But when the researchers restricted the view to pages published after ChatGPT's late-2022 release, the share jumped to over a third. That single contrast is the key to the whole confusion: the same dataset produces 10% or over a third depending only on whether you include the old, pre-AI web in the denominator. The web is full of pages that predate ChatGPT entirely, and every one of them drags the all-pages average down. Filter to recent pages and the rate more than triples.
Pew is careful about interpretation, and so should anyone citing it be. The figure is not a claim that a third of the internet was written start-to-finish by a machine. It measures pages that show signs AI played a role in writing or heavily editing them — assistance and editing count, not just full generation. And detection is imperfect: like every AI detector, Open Pangram misclassifies some human writing as AI and some AI writing as human. The defense is scale. Across hundreds of thousands of documents the aggregate pattern is reliable even though any single page's label may be wrong, so the honest reading is 'directionally accurate,' not 'a precise verdict on this page.'
If you take one thing from the Pew study, make it the distribution rather than the average. The rate of AI authorship varied sharply by domain type, and the variation is far more actionable than the single headline number. About 10% of .com pages carried AI signals — roughly double the 4.6% share on .org pages, and about ten times the rate on .edu and .gov pages, which both sat near 1%. AI writing is not spread evenly across the web; it is concentrated where content is produced at commercial volume, and it is scarce where institutional trust and verification are the whole point.
This tells you where you actually compete. If you publish a business blog, a marketing site, or SEO content, you are publishing into the .com tier — the exact part of the web where AI text is densest and the reader's guard is highest. The near-1% rate on .edu and .gov is not a quirk; it is a signal about what earns trust. Where being wrong has consequences and where a human institution stakes its name on the page, human-sounding, verifiable writing still dominates. That is the bar a serious brand should hold itself to, and it reframes the goal: not 'avoid looking like AI on average,' but 'read like the high-trust corners of the web, in a space where almost nobody else does.'
Graphite, a growth and SEO firm, produced the other widely-cited 2026 figure by sampling Common Crawl and classifying articles with AI detectors. Its finding: among newly published English-language articles, the share that are primarily AI-generated rose from a near-zero baseline before ChatGPT to roughly parity with human-written articles by early 2025, and has plateaued near half since. A follow-up analysis that averaged three detectors — Pangram, Copyleaks, and GPTZero — across about 55,000 articles into early 2026 put the figure near 49.9%, barely moved from the single-detector version, which is a sign the trend is robust.
Line Graphite's method up against Pew's and the gap explains itself. Graphite counts only newly published articles, not all pages — so it excludes the entire pre-AI web that pulls Pew's all-pages number down to 10%. It counts articles specifically, not every page type. And its threshold is 'primarily AI-generated,' meaning more than half the text reads machine-written, which is a different bar from Pew's broader 'shows signs of AI authorship.' Put those together and the comparison to make is not Pew's 10% against Graphite's 50% — it is Pew's over-a-third figure for recent pages against Graphite's ~50% for recent articles. Aligned on population, the two studies are far closer than the headlines suggest, and they agree on the shape of the curve: a fast climb after ChatGPT, then a plateau.
Because these numbers travel fast and lose their context, it is worth having a fixed checklist. Before you repeat or act on any 'this share of the web is AI' figure, ask four things — they account for essentially all of the variance between honest studies.
All pages ever crawled, or only recently published ones? Articles, or every page type? This is the biggest lever by far. The pre-AI web is enormous and permanent, so any 'all pages' figure will always run lower than any 'new content' figure. Pew's 10% and over-a-third figures come from the same data and differ only on this axis. When a number sounds surprisingly high or low, the population is usually why.
Does the study count any detectable AI involvement, or only content that is majority machine-written? 'Shows signs of AI' (Pew) sweeps in AI-assisted and heavily-edited human work; 'primarily AI-generated' (Graphite) requires most of the text to read as machine output. A looser threshold on a broader population and a stricter threshold on a narrower one can produce numbers that look contradictory while describing the same web.
One model or several averaged together? Every AI detector is probabilistic and carries its own biases, false positives, and false negatives. A single-detector figure inherits that tool's error; averaging independent detectors, as Graphite's follow-up did, reduces it. No detector is authoritative, so a figure built on one should be read with wider error bars than one built on several.
On a curve that rose from near zero to near parity in about two years, the snapshot date matters. A figure from early 2025 and one from mid-2026 can differ simply because the underlying share moved, not because the studies disagree. Always note the measurement date, and be suspicious of any stat quoted without one.
Strip away the number and the same conclusion survives every version of the study. A large and growing share of the web — most of the newly published commercial web — now reads as machine-written, carrying the same handful of statistical tells a detector is trained to count. Pew catalogued those tells: since 2023 the frequency of em dashes has roughly doubled, Oxford commas are up about 63%, words like 'delve,' 'interplay,' and 'testament' have surged, and 'negative parallelism' — the 'it's not X, it's Y' construction — has nearly tripled. When a third of new pages share that fingerprint, writing that matches the fingerprint is writing no reader and no ranking system has a reason to prefer.
This connects to a finding that outlives the exact percentages: volume and distribution have come apart. The web filled with AI text, but that text mostly is not the content winning search visibility or getting cited by AI assistants — a decoupling covered in depth in the guide on AI-generated content flooding every platform and, on the search side, scaled AI content and crawl economics. Producing more of the commodity everyone is producing does not buy reach; it buys a place in the pile these studies measure. The statistic is not a scoreboard you climb by generating more. It is a description of what everyone else's output now looks like — which makes it a map of where the openings are.
The tempting misread is to treat this as a detection game — run your content through a checker, tune it until the score says 'human,' and publish. This fails for two reasons. First, the detectors are noisy in both directions, so a clean score is not proof of anything, and the tools that promise to 'humanize' AI text are themselves increasingly caught. Second, and more fundamentally, gaming a detector optimizes for the wrong target. The goal is not to pass a scan; it is to be genuinely worth a reader's time and a platform's trust. A page that is technically undetectable but says nothing original still reads like the web's new average, and both readers and ranking systems discount it. Detection is a symptom you can chase forever; the disease is sameness, and you treat sameness with substance, not with a rewriter. The mechanics of why detectors misfire are covered in the guide on AI content detection.
Three things reliably put content on the side of the line these studies are not measuring. The first is a governed identity — a specific, recorded point of view and voice, so output reads as a particular someone rather than as generic model prose. Generic AI writing sounds generic because it was never told to be anything in particular; give the model a real specification of who it is writing as and the fingerprint changes. The second is format. The studies count text, and the surge is densest in the cheapest text formats; short-form video, avatar content, brand-exact carousels, and native social posts are far harder to mass-produce from a one-line prompt and compete in feeds that are not yet the article graveyard. The third is provenance — a real human accountable for the result, which is both what the high-trust .edu/.gov corners have and what platform enforcement increasingly demands.
Notice these are not detector tricks; they are properties of the work itself. That is the point. You cannot out-measure or out-produce a web that is already a third machine-written on recent pages. You can only make sure your own output is the exception — the kind that carries an identity a model cannot average away, in formats a text scan never sees, with a person who stands behind it.
Kompozy is an AI content generation and multi-platform publishing engine, and the honest way to place it against these numbers is by which surface it touches. Only one slice of what it produces — its blog and article generation, publishing to your own destinations — lands on the crawled, indexable web these studies measure. That is the surface where the .com AI-density problem is real, so it is the surface that gets the most editorial discipline: blog output is generated from a written Persona Brief that encodes a real voice and point of view, filtered against the exact clichés Pew catalogued — the reflexive 'delve,' the negative-parallelism template — and passed through a per-post human review gate before anything ships. That is a workflow of fewer, governed, reviewed pieces with a person accountable for each one: the provenance the high-trust web has, and the inverse of the scaled fingerprint the detectors count.
The larger share of what Kompozy makes lives off the crawled web entirely, which is precisely where the format lever points. From one source idea it generates captioned Persona Shorts and longer avatar video, brand-exact Carousel Posts rendered through HyperFrames, quote graphics, photo posts, and email newsletters — native assets a text detector never looks at, each carrying a consistent persona face and voice a model cannot average away. An AI Influencer persona pool keeps that identity stable across everything, and Autopilot schedules and fans it across the eight social platforms plus blog and email, with the review pipeline in front of publishing. The web average is not the number you fix; being identifiably yourself, in formats and channels the flood has not reached, is — and that is the shape the tool is built around.
Ten percent and fifty percent are both correct answers to 'how much of the web is AI-generated,' because they answer different questions: Pew counts all pages and finds about a tenth carry AI signals, over a third among recent ones, concentrated on commercial .com sites; Graphite counts new articles and finds roughly half. Reconciled on population, threshold, detector, and date, the studies agree — AI text is a large, fast-growing, commercially-concentrated share of the web that has plateaued near parity for fresh content. The practical conclusion does not depend on which figure you trust. The web now has a statistical average voice, and blending into it is the one strategy guaranteed to fail. The durable move is to be the exception these numbers are not counting: a recognizable identity, formats a text scan never sees, and a human standing behind the work — the profile that still earns trust in the one corner of the web AI has barely touched.
It depends entirely on what you count. Pew Research's 2026 analysis of nearly 490,000 Common Crawl pages found about 10% of all pages in a July 2026 sample show signs of AI authorship, rising to over a third among pages published after ChatGPT's late-2022 release. Graphite, sampling only newly published articles rather than all pages, put the primarily-AI share near half. Both are defensible; they measure different populations, so cite the one that matches your question — all pages, recent pages, or new articles.
Four things move the figure. The population — all pages ever crawled versus only recently published articles — is the biggest lever, because old pages and non-article pages dilute the share. The threshold — 'shows any sign of AI' versus 'more than half the text reads machine-written' — sets a different bar. The detector — one model versus an average of several — changes the error. And recency — a 2025 snapshot versus a 2026 one — matters on a fast-rising curve. Pew and Graphite differ on all four, which is why 10% and 50% can both be right.
Published August 20, 2026, Pew ran nearly 490,000 English-language Common Crawl pages (January 2021–July 2026) through the Open Pangram detection model. In a random July 2026 sample, about 10% of pages overall showed AI-authorship signals, and over a third of pages published after ChatGPT's release did. The rate concentrated on commercial sites: about 10% of .com pages versus 4.6% of .org and near 1% on .edu and .gov. Pew stresses the figure is directionally reliable at scale, not a precise per-page verdict.
Commercial ones. In Pew's data, about 10% of .com pages showed signs of AI authorship — roughly double the 4.6% on .org and about ten times the near-1% rate on .edu and .gov. AI writing concentrates where content is produced at volume for commercial ends and stays rarest where institutional trust and verification matter most. That distribution is more useful than the headline average: it tells you the AI-text problem is worst in exactly the commercial, SEO-driven space most creators publish into.
Reliable as directional estimates, not as a census. Both studies sample Common Crawl, which is broad but not the whole web, and both lean on AI detectors, which are probabilistic and misclassify individual pages in both directions. The aggregate across hundreds of thousands of documents is trustworthy even when any single label is wrong; the exact percentage is not. Read every such figure as 'roughly this share of this population,' and check the population, threshold, detector, and date before repeating it.
Both are right because they count different things. Pew's 2026 analysis of nearly 490,000 Common Crawl pages found roughly 10% of all pages — and over a third of pages published since ChatGPT's launch — show signs of AI authorship. Graphite, sampling only newly published articles, put the AI share near half. Different populations, detectors, and thresholds produce different numbers, but both point the same way: AI text is now a large, fast-growing slice of the web, concentrated on commercial sites.
Get started → · ← All guides · Compare Kompozy vs other tools