# What percent of webpages will be significantly written or edited by AI in 2030?

Canonical URL: https://preseen.com/reports/38912575-44e6-4e5d-a8e7-6a1eb5cfed03/what-percent-of-webpages-will-be-significantly-written-or-edited-by-ai-in-2030
Markdown URL: https://preseen.com/reports/38912575-44e6-4e5d-a8e7-6a1eb5cfed03/markdown

## Forecast

Median forecast: 28; 80% interval: 17.5 to 45.

Generated: September 2, 2026 at 12:03 AM UTC
Forecast model: gpt-5.6-sol
Research model: gpt-5.6-sol

## Analysis

## TL;DR

My conditional forecast centers on **30%** of English-language webpages showing significant AI authorship or editing in the qualifying 2030 snapshot. The median is 28%, with an 80% interval from 17% to 44.7% and a 6% chance of exceeding 50%. The old human-written stock slows the result, while AI-heavy new publishing and AI rewrites keep the upper tail large.

## Context

[Pew's August 20, 2026 study](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) is the direct baseline. Its July 2026 Common Crawl sample found significant AI signals on 10% of all pages; contemporaneous coverage reported the chart value as 9.6%. Pew sampled 10,000 English-language pages from each of 49 crawls from January 2021 through July 2026, for N=490,000, and classified Open Pangram scores of at least 0.2 as meaningful AI authorship or editing ([methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

The target is a stock measure. Old pages remain in the denominator. I model the latest likely 2030 snapshot, roughly 4.3 years after July 2026, rather than treating the share of newly published articles as the answer.

## Evidence

The historical backbone is Pew's crawl series. Pew does not provide the all-page monthly values as a readable table, so I built a consistent proxy by weighting its complete `.com`, `.org`, `.edu` and `.gov` table by the reported sample counts. The first 11 rows below are that four-domain proxy; the final row is the actual all-page headline. The proxy omits other domains, so its level is approximate, but its trend is useful.

| Period | Four-domain weighted rate | N across four domains | Basis |
|---|---:|---:|---|
| 2021 H1 | 1.02% | 36,654 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2021 H2 | 0.98% | 29,150 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2022 H1 | 0.98% | 14,378 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2022 H2 | 0.97% | 29,941 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2023 H1 | 1.46% | 23,267 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2023 H2 | 2.51% | 14,875 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2024 H1 | 3.43% | 28,837 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2024 H2 | 4.21% | 43,159 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2025 H1 | 4.96% | 42,873 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2025 H2 | 5.96% | 41,536 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| 2026 H1 | 8.36% | 49,719 | [Pew data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |
| July 2026 | 10.0% all-page headline | 10,000 | [Pew headline](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/) |

Simple extrapolation gives a wide model range. A linear fit using all post-2022 points reaches 18.7% in late 2030. A fit using only the faster mid-2024 onward trend reaches about 22%. Flexible quadratic and Gompertz fits reach roughly 28% to 30%. I treat the first pair as low cases because they do not model continued diffusion into editing workflows; the flexible curves are too sensitive to four years of accelerating data to stand alone.

New content is already much more AI-heavy than the stock. A study by researchers at Imperial College London, the Internet Archive and Stanford sampled 33 monthly publication cohorts from August 2022 through May 2025 and found that roughly 35% of newly published websites were AI-generated or AI-assisted by mid-2025 ([Dolezal et al.](https://arxiv.org/abs/2604.26965)). Pew found the same 35% rate among the non-random subset of July 2026 pages with detectable post-ChatGPT publication dates; only 10% to 15% of sampled pages had usable dates ([Pew methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

Article-only estimates are higher. Graphite randomly selected 55,400 English articles and listicles from Common Crawl, dated January 2020 through March 2026, and averaged Pangram, Copyleaks and GPTZero classifications. Primarily AI-generated articles reached roughly 50% by early 2025 and remained near that level for five quarters through the first quarter of 2026 ([Graphite's May 2026 update](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do)). Ahrefs classified 74.2% of 900,000 newly detected English pages from April 2025 as containing some AI material, but only 2.5% as pure AI; its broad proprietary definition makes this an upper-bound signal, not a substitute for Pew's threshold ([Ahrefs study](https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated/)).

The other key variable is turnover. Pew sampled nearly one million Common Crawl pages from 2013 through 2023 and found that 38% of the 2013 cohort was inaccessible by October 2023, while about 20% of the 2021 cohort was gone after two years ([Pew link-rot study](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/)). Those figures correspond to rough annual disappearance rates of 4.7% and 10.6%. Substantial edits add another conversion channel, but studies that count every checksum or template change overstate the rate relevant to AI authorship.

Common Crawl snapshots also change faster than webpage deaths alone imply. The weighted share of unique URLs that had never appeared in any earlier Common Crawl archive was 43.2% in 2021, 43.5% in 2022, 38.7% in 2023, 37.8% in 2024, 31.5% in 2025 and 28.4% through the first six crawls of 2026, based on my calculations from Common Crawl's [monthly URL totals](https://raw.githubusercontent.com/commoncrawl/cc-crawl-statistics/master/plots/crawlsize/monthly.csv) and [new-URL estimates](https://raw.githubusercontent.com/commoncrawl/cc-crawl-statistics/master/plots/crawlsize/monthly_new.csv). These are URLs new to the archive, not necessarily newly published pages. They show why a crawl snapshot is neither a fixed stock nor a pure sample of new content.

I use the stock-flow equation

$$
\frac{dS}{dt}=r(t)\bigl(F(t)-S(t)\bigr),
$$

where \(S\) is the flagged share of the crawl, \(F\) is the flagged share among new or substantially rewritten pages, and \(r\) is the effective annual rate of replacement, substantive editing and crawl recomposition. Matching the rise from the roughly 1% pre-ChatGPT detector floor to 9.6% in July 2026, while respecting the 35% to 50% flow evidence, implies a historical effective rate near 9% to 13% per year. Rates above 20% would generally have produced a higher 2026 stock unless the true broad-page flow was far below the available estimates.

For 2030, my central path has the broad-page flow rising from roughly 40% to 45% in 2026 to 60% to 65%, while the effective conversion rate rises modestly as AI-assisted rewriting becomes easier. That produces a stock result in the high 20s. The lower path assumes the flow plateaus and durable pages dominate. The upper path assumes AI becomes the standard drafting layer, existing URLs are rewritten faster, and the future study detects mixed authorship better.

Google supplies a real brake, but not a ban. Its policy targets large-scale, low-value production regardless of whether humans or machines produced it, and its March 2024 changes cut low-quality, unoriginal search results by a reported 45% ([Google](https://blog.google/products-and-platforms/products/search/google-search-update-march-2024/)). Graphite found that AI-generated articles were about half of new article production but only 14% of articles appearing in Google results ([Graphite search study](https://graphite.io/five-percent/research/ai-content-in-search-and-llms)). Common Crawl prioritizes pages using host budgets and link-based harmonic centrality, so weak visibility can also reduce crawl representation, though Google rank and Common Crawl selection are not the same process ([Common Crawl analysis](https://arxiv.org/abs/2607.13636)).

Measurement creates the widest uncertainty. Pew's open and commercial Pangram models agreed on 96% of 62,370 pages, but Cohen's kappa was only 0.61, and the open model showed an approximately 1% pre-ChatGPT false-positive floor ([Pew methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)). The six-million-generation RAID benchmark found large failures on unseen generators, changed decoding and adversarial rewrites ([RAID](https://arxiv.org/abs/2405.07940)). Mixed human-machine writing is harder still: the MixSet experiments found clear limits in detectors applied to AI-revised human text and human-revised AI text ([NAACL study](https://aclanthology.org/2024.findings-naacl.29/)).

I encode that evidence in four regimes:

| Regime | Weight | Center | SD | Main interpretation |
|---|---:|---:|---:|---|
| Slow turnover or flow plateau | 25% | 20% | 5 points | Durable legacy stock and strong pruning of low-value AI pages |
| Central diffusion | 55% | 28.5% | 7 points | Moderate flow growth and historically calibrated turnover |
| Faster adoption and rewriting | 17% | 42.0% | 10 points | AI becomes standard across publishing and editing workflows |
| AI-default or broader measurement | 3% | 60.0% | 12 points | Autonomous publishing or a much more sensitive mixed-authorship measure |

The bounded mixture has a mean of 30% and a median of 28%. Its central 50% interval is 22% to 35.2%, its central 80% interval is 17% to 44.7%, and its central 90% interval is 15% to 51.9%. It assigns 6% probability above 50% and less than 1% below 10%.

## What's non-obvious

The web-stock argument is often made too mechanically. Using URL deaths alone understates refresh because Common Crawl continually discovers pages and revisits changing sites. Treating every new-to-crawl URL as freshly published makes the opposite error, because an old human-written page can be new to the archive. The observed 2022–2026 rise is the best calibration of the combined process.

The apparent 50% plateau in AI-written articles is not a plateau in the question's target. Graphite asks whether an article is primarily AI-generated. Pew also counts substantial AI editing and covers product pages, company pages, reference pages and other non-article text. A future dominated by human-reviewed AI drafts can therefore raise Pew's measure while the pure-AI article share stays flat.

## Uncertainties

- No source provides a reliable age histogram for all English pages in each Common Crawl snapshot. That would sharply improve the stock-flow estimate.
- No large panel follows the same URLs over time and records when substantive text revisions change their AI classification. Existing change studies mix real rewriting with ads, dates, counters and templates.
- Detector performance against 2028–2030 models and ordinary human editing cannot be tested today. Conceptual comparability reduces this problem but does not remove it.
- Common Crawl's domain mix, robots exclusions and crawl priorities may change before 2030. Those choices can move the reported percentage even if underlying author behavior changes smoothly.
- The numeric distribution is conditional on a qualifying study being published. The resolution rules permit annulment, but the required 202-element output has no annulment bucket, so that probability is not included in the array.

## Sources

- Domain Expert Search (mcp)
  > Found 14 domain experts for 'forecasting AI-generated and AI-edited text prevalence on the public web using web-crawl sampling, stock-flow dynamics, detector validity, and content turnover':
- Claude Code (e2b)
  > Job coding_whiz_job_eaab76b867 done after 799810ms.
- [pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai) (tool)
- [pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content) (tool)
- [arxiv.org](https://arxiv.org/abs/2604.26965) (tool)
- [originality.ai](https://originality.ai/ai-content-in-google-search-results) (tool)
- [newsguardtech.com](https://www.newsguardtech.com/special-reports/ai-tracking-center) (tool)
- [pewresearch.org](https://www.pewresearch.org/internet/2024/05/17/when-online-content-disappears) (tool)
- [dl.acm.org](https://dl.acm.org/doi/10.1145/988672.988674) (tool)
- Domain Expert Research Task (mcp)
  > Job domain_expert_research_task_1600995161 done after 316859ms.
- [the-decoder.com](https://the-decoder.com/pew-study-shows-ai-written-text-has-surged-across-the-web-since-late-2022) (tool)
- [noah-news.com](https://noah-news.com/pews-35-percent-ai-web-finding-comes-from-pages-that-publish-a-date) (tool)
- [thenextweb.com](https://thenextweb.com/news/pew-ai-authorship-web-pages-study) (tool)
- [pangram.com](https://www.pangram.com/blog/introducing-open-pangram) (tool)
- [graphite.io](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do) (tool)
- [ai-on-the-internet.github.io](https://ai-on-the-internet.github.io/) (tool)
- [gptzero.me](https://gptzero.me/news/ai-written-web) (tool)
- [ahrefs.com](https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated) (tool)
- [Link Rot and Digital Decay on Government, News and Other Webpages | Pew Research Center](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears) (openai)
- [pewresearch.org](https://www.pewresearch.org/wp-content/uploads/sites/20/2024/05/pl_2024.05.17_link-rot_report.pdf) (tool)
- [zenodo.org](https://zenodo.org/records/13687116) (tool)
- [ws-dl.blogspot.com](https://ws-dl.blogspot.com/2024/09/2024-09-20-some-urls-are-immortal-most.html) (tool)
- [doi.org](https://doi.org/10.1002/asi.10018) (tool)
- [doi.org](https://doi.org/10.1002/(sici) (tool)
- [doi.org](https://doi.org/10.1002/asi.23561) (tool)
- [i.stanford.edu](http://i.stanford.edu/~olston/publications/webstudy.pdf) (tool)
- [marc.najork.org](https://marc.najork.org/papers/spe2004.pdf) (tool)
- [gcn.com](https://gcn.com/ai-authorship-one-ten-webpages-now/21102) (tool)
- [webiano.digital](https://webiano.digital/the-ai-text-crossover-will-arrive-sooner-in-chat-than-on-the-open-web) (tool)
- [searchenginejournal.com](https://www.searchenginejournal.com/1-in-10-webpages-shows-signs-of-ai-authorship-pew-reports/586883) (tool)
- [isoclive.substack.com](https://isoclive.substack.com/p/un-ai-report) (tool)
- [aibizinsider.com](https://aibizinsider.com/2026/08/21/ai-authored-web-pew-study-2026-08-21) (tool)
- [hdl.handle.net](http://hdl.handle.net/10451/14153) (tool)
- [report-ai.org](https://report-ai.org/indexes/enterprise-ai/ai-generated-content-statistics-2026) (tool)
- [ahrefs.com](https://ahrefs.com/blog/google-doesnt-punish-ai-content) (tool)
- [serpary.com](https://serpary.com/en/news/ahrefs-331k-pages-google-ai-content-quality-study) (tool)
- [arxiv.org](https://arxiv.org/abs/2605.00087) (tool)
- [developers.google.com](https://developers.google.com/search/blog/2023/02/google-search-and-ai-content) (tool)
- [blog.google](https://blog.google/products-and-platforms/products/search/google-search-update-march-2024) (tool)
- [buttonblock.com](https://buttonblock.com/blog/google-spam-policy-ai-generated-content-2026) (tool)
- [elevarus.com](https://elevarus.com/does-google-august-2026-spam-update-penalize-ai-content) (tool)
- [searchenginejournal.com](https://www.searchenginejournal.com/reports-indicate-googles-spam-update-focused-on-seo-ai-content/586978) (tool)
- [designcopy.net](https://designcopy.net/en/scaled-content-abuse-ai-publishing-2026) (tool)
- [poweredby.keywordseverywhere.com](https://poweredby.keywordseverywhere.com/ai/ccbot) (tool)
- [ustechautomations.com](https://ustechautomations.com/resources/blog/how-many-sites-block-ccbot-2026) (tool)
- [cloudflare.com](https://www.cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large) (tool)
- [wpnews.pro](https://wpnews.pro/news/common-crawl-published-a-manual-for-being-visible-to-ai-i-automated-it) (tool)
- [w3techs.com](https://w3techs.com/technologies/details/cm-wordpress) (tool)
- [wordpress.org](https://wordpress.org/plugins/wds-mcp-content-manager) (tool)
- [storeinspect.com](https://storeinspect.com/blog/shopify-ai-app-adoption) (tool)

## Question Details

This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. (pewresearch.org) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.

### Resolution Criteria

Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. (pewresearch.org) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. (pewresearch.org) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.

### Fine Print

“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. (pewresearch.org) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.
