# What percent of webpages will be significantly written or edited by AI in 2030?

Canonical URL: https://preseen.com/reports/5de82363-a9f7-4bc5-98d5-4e2abb39d7fe/what-percent-of-webpages-will-be-significantly-written-or-edited-by-ai-in-2030
Markdown URL: https://preseen.com/reports/5de82363-a9f7-4bc5-98d5-4e2abb39d7fe/markdown

## Forecast

Median forecast: 27; 80% interval: 14.5 to 49.

Generated: October 7, 2026 at 11:35 AM UTC
Forecast model: gpt-6.1-sol
Research model: gpt-5.6-sol+gpt-6.1-sol+claude-opus-5-5

## Analysis

## TL;DR
My point forecast is **27%**, with a distribution mean of 29% ([forecast calculation](#evidence)). The central 80% predictive interval is 14%–48.6%, and the probability of exceeding 50% is 9% ([forecast calculation](#evidence)). These probabilities concern all sampled English-language webpages and are conditional on a qualifying numerical resolution, not annulment ([resolution criteria](https://preseen.com/questions/0cf01fa6-771a-4e23-8cf6-4bfcff7878a1)).

## Context
The target is a page-level classification of meaningful AI writing or substantial editing. It is not the share of AI-generated words, newly published articles, websites, or page views. The starting observation is 9.60% of all sampled pages in July 2026, reported as 10% in Pew’s August 20, 2026 release ([Pew study](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)). Pew sampled 10,000 English-language pages from each of 49 crawls, January 2021–July 2026, and classified extracted text with Open Pangram at a score of at least 0.2 ([methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

The forecast uses the client’s October 7, 2026, 11:21 UTC cutoff and targets a snapshot collected during 2030. Earlier observations are forecasting evidence, not resolution events. The latest qualifying 2030 snapshot takes precedence; failure to obtain a qualifying study by December 31, 2032 means annulment rather than a numerical outcome ([client’s resolution criteria](https://preseen.com/questions/0cf01fa6-771a-4e23-8cf6-4bfcff7878a1)).

## Evidence
The historical backbone is the complete all-page series below. Entries give month-day and detector-positive percentage within each year, preserving the chart’s displayed precision. These are all 49 observations from the August 20 study, as available at the cutoff, with 10,000 sampled page observations per point ([Pew chart data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/); [sampling design](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

| Year | Crawl date: flagged percentage |
|---|---|
| 2021 | 01-27:1.18; 03-03:1.13; 04-20:1.10; 05-17:1.12; 06-23:1.08; 08-01:1.11; 09-19:1.07; 10-20:1.13; 11-30:1.09 |
| 2022 | 01-24:1.10; 05-25:1.04; 07-06:1.04; 08-16:1.02; 09-29:1.07; 12-08:1.15 |
| 2023 | 02-04:1.35; 03-20:1.53; 06-09:1.89; 09-21:2.30; 12-10:2.78 |
| 2024 | 03-01:3.22; 04-13:3.55; 05-18:3.49; 06-24:3.49; 07-24:3.40; 08-04:3.80; 09-20:3.83; 10-12:4.34; 11-02:4.51; 12-10:4.92 |
| 2025 | 01-22:4.86; 02-06:4.92; 03-23:4.98; 04-22:5.07; 05-12:4.93; 06-22:4.86; 07-14:4.78; 08-13:5.24; 09-12:5.49; 10-15:6.01; 11-15:6.29; 12-11:6.64 |
| 2026 | 01-25:6.92; 02-19:7.28; 03-15:8.09; 04-20:8.58; 05-19:8.86; 06-06:8.98; 07-18:9.60 |

Growth has accelerated, but it has also paused. My linear fits projected to December 31, 2030 produce the following results; these are calculations from the complete series, not Pew forecasts ([source data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)).

| Fitting window | Observations | Slope, percentage points/year | Projected share |
|---|---:|---:|---:|
| December 8, 2022–July 18, 2026 | 35 | 2.11 | 18% |
| March 1, 2024–July 18, 2026 | 29 | 2.46 | 19% |
| January 22, 2025–July 18, 2026 | 19 | 3.36 | 24% |
| January 25–July 18, 2026 | 7 | 5.55 | 34.4% |

The fitting window matters more than ordinary sampling error. Bounded logistic fits to the same post-launch history also fit similarly well while producing very different endpoints when their assumed ceilings change. I therefore treat both recent acceleration and slower long-run growth as evidence about the range, not as grounds for a narrow extrapolation ([underlying series](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)).

Legacy pages are the main restraint. Pew’s May 17, 2024 persistence study sampled just under one million URLs from capture-year cohorts spanning 2013–2023 and checked accessibility in fall 2023. It found that 38% of the oldest cohort had disappeared, leaving most URLs accessible after roughly a decade. That measures URL survival, not unchanged text. It supports persistent legacy content but does not identify the rate at which surviving pages are substantially rewritten ([persistence study](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/)).

Incoming prose is much more AI-intensive than the whole stock. Graphite’s May 15, 2026 study analyzed 55,400 eligible English articles published from January 2020 through March 2026. Its latest publication cohort was roughly half primarily AI-generated, and the authors reported a plateau. Eligibility required article markup, a publication date, and at least 100 words. Its detector validation did not cover human-in-the-loop rewriting. I use it to support high incoming adoption, not as an all-page baseline or a precise ceiling ([Graphite study](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do)).

Existing pages can convert without disappearing. WordPress.com’s March 20, 2026 announcement, updated May 15, describes agents creating and updating posts and pages on paid plans. This establishes an available mechanism, not a measured adoption rate. I read it as upward pressure on substantial legacy-page editing ([product announcement](https://wordpress.com/blog/2026/03/20/ai-agent-manage-content/)). Search incentives push the other way: Google permits useful AI-assisted content but warns against scaled generation without added value. That limits some publishing incentives without removing every such page from the public web ([Google guidance](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content)).

Measurement changes can move the headline in either direction. Pew’s comparison on 62,370 pages found 96% agreement between its open model and Pangram 3.3, but agreement is not ground-truth accuracy ([Pew validation](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)). In Pangram’s July 29, 2026 vendor evaluation, 14,990 substantially edited student texts were classified Mixed-or-AI at rates of 21% by Pangram 3.3.2 and 58.6% by Pangram 4. Those are results on the same selected benchmark, not a correction to web prevalence ([technical report, Table 5](https://arxiv.org/html/2607.27183v1)).

Independent evidence argues against assuming detection inevitably collapses. Epoch’s July 15, 2026 evaluation, using June detector versions, found no Pangram errors among 495 human passages or 297 ordinary-prompt AI passages. Style imitation produced more misses; counting Mixed as detected gave a 6% miss rate. These were roughly 500-word prose passages, not representative webpages ([Epoch evaluation](https://epoch.ai/data-insights/ai-detectors-false-negatives)). Research published June 23 also found that adaptive methods could recover detection under shifts that defeated fixed detectors, although its experiments used scientific abstracts and benchmark text rather than a random web sample ([distribution-shift study](https://arxiv.org/html/2606.25152v1)). I widen both tails instead of applying a blanket detection penalty.

My quantitative core is a stock-and-editing model:

$$
\frac{dp}{dt}=r\bigl(q(t)-p(t)\bigr)+e\bigl(1-p(t)\bigr).
$$

Here, \(p\) is the share positive on a reference classification scale; \(r\) is the annual effective renewal hazard from incoming pages, growth-induced dilution, and displacement; \(q\) is the positive share of incoming material; and \(e\) is an additional substantial-editing hazard among remaining nonpositive pages. This is not a claim to know true provenance. I start at \(p=0.096\) and use a 4.25-year central horizon, with snapshot timing uncertainty included in the scenario spreads ([baseline](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/); [forecast calculation](#evidence)).

The main scenario uses \(r=0.10\) and \(e=0.015\) per year, with \(q\) rising linearly from 0.40 to 0.60. It initially adds 4.4 percentage points per year and ends near 28%. Slower and faster parameter sets preserve disagreement about renewal, editing, and incoming adoption. All rates, weights, and spreads are my judgments, not measured population parameters ([forecast calculation](#evidence)).

| Scenario | Weight | Center, percent | Within-scenario SD, percentage points |
|---|---:|---:|---:|
| Slow growth plus strong measurement drag | 5% | 10 | 5 |
| Persistent legacy stock and slow diffusion | 22% | 18 | 5 |
| Continued mainstream diffusion | 53.0% | 28 | 7 |
| Rapid renewal, editing, and improved detection | 16% | 45.3 | 10 |
| Large-scale automated page proliferation | 4% | 72 | 12 |

Centers are rounded for display. The first is a separate measurement stress case; the others come from the stock equation. The mixture keeps a right tail without treating near-total AI dominance as the default ([forecast calculation](#evidence)).

I represent each scenario with a bounded Beta distribution and integrate prospective sampling noise through a Beta-binomial model, assuming 10,000 pages. Reporting assumptions allocate 70% to an available unrounded estimate, 10% to one-decimal reporting, and 20% to a whole-percent headline. Rounded outcomes enter their correct buckets rather than being spread across adjacent buckets. The resulting distribution assigns 4% to results at or below 10% and 9% to results above 50%; numerical precision in the code is not precision in the underlying assumptions ([forecast calculation](#evidence)).

## What's non-obvious
**Crawler novelty is not content renewal.** Common Crawl’s July 28, 2026 release describes 2.14 billion captures collected July 7–25, including 603 million URLs never previously visited. Those URLs can contain old text. Treating that count as monthly replacement would vastly overstate how quickly AI-intensive new writing displaces the legacy stock ([crawl release](https://commoncrawl.org/blog/july-2026-crawl-archive-now-available)). Likewise, the dated subset is nonrandom, so its AI share cannot identify the age composition of the whole web ([Pew methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

Text filtering can also create apparently conflicting estimates. A September 30, 2026 preprint audited 10,000 raw documents from the July crawl. FineWeb retained 282 of 963 AI-labeled documents, versus 1,104 of 8,649 human-labeled documents—about 2.3 times the retention rate. The audit included language filtering, and the labels were detector judgments. This demonstrates selection bias, not a population correction. Filtered token shares and unfiltered page shares should not be averaged together ([filter audit, Appendix A.1](https://arxiv.org/html/2609.40295v1)).

## Uncertainties
The largest gaps concern the page stock, not generation cost:

- A representative distribution of page ages and substantive revision rates would identify renewal and editing separately. URL-survival evidence does neither ([persistence study](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/)).
- Validation on short, templated, mixed-author, and substantially edited webpages would sharpen measurement uncertainty. Current prose benchmarks do not supply that population-weighted sensitivity ([Pangram evaluation](https://arxiv.org/html/2607.27183v1); [Epoch limitations](https://epoch.ai/data-insights/ai-detectors-false-negatives)).
- Changes in crawler access, extraction, and filtering can alter the denominator. Research documents restrictions on Common Crawl’s crawler, but does not establish their future net effect on AI prevalence ([crawler study](https://cseweb.ucsd.edu/~savage/papers/IMC25Crawlers.pdf)).

Methodological comparability remains a real boundary. An October 5, 2026 paper separates AI-origin ideas from AI-written prose; idea-only involvement is not automatically the target here ([IdeaLens](https://arxiv.org/abs/2610.06778)). A future study must still measure substantial text authorship or editing. The forecast does not assign annulment to zero, and it does not extrapolate another year into a numerical resolution when the required study is missing ([resolution criteria](https://preseen.com/questions/0cf01fa6-771a-4e23-8cf6-4bfcff7878a1)).

## Sources

- Domain Expert Search (mcp)
  > Found 6 domain experts for 'AI text detection research Pangram mixed human AI edited text detector validity and aggregate web prevalence measurement':
- arXiv (mcp)
  > Found 10 total papers. Showing 10:
- [github.com](https://github.com/RishanthRajendhran/IdeaLens) (tool)
- [huggingface.co](https://huggingface.co/collections/rishanthrajendhran/idealens-6abee785ce6196fc0be9200f) (tool)
- [ideadetector.ai](http://ideadetector.ai/) (tool)
- [arxiv.org](https://arxiv.org/pdf/2610.06778v1) (tool)
- [arxiv.org](https://arxiv.org/pdf/2606.25152v1) (tool)
- [github.com](https://github.com/kkr36/llm_detection) (tool)
- Claude Code (e2b)
  > Job coding_whiz_job_e31afc0b18 done after 284407ms.
- [How Much of the Internet Is Written With AI? | Pew Research Center](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai) (openai)
- [Methodology | Pew Research Center](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content) (openai)
- Domain Expert Research Task (mcp)
  > Job domain_expert_research_task_b70a1c0edf done after 413573ms.
- [Pangram 4 Technical Report](https://arxiv.org/html/2607.27183v1) (openai)
- [AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors | Epoch AI](https://epoch.ai/data-insights/ai-detectors-false-negatives) (openai)
- [Agents Can Use Base Models to Evade AI Detection](https://arxiv.org/html/2609.31876v1) (openai)
- [How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text](https://arxiv.org/html/2609.40295v1) (openai)
- [huggingface.co](https://huggingface.co/pangram/editlens_Llama-3.2-3B) (tool)
- [arxiv.org](https://arxiv.org/abs/2604.26965) (tool)
- [The Impact of AI-Generated Text on the Internet](https://arxiv.org/html/2604.26965v1) (openai)
- [arxiv.org](https://arxiv.org/html/2412.18148v3) (tool)
- [arxiv.org](https://arxiv.org/abs/2510.07226) (tool)
- [arxiv.org](https://arxiv.org/html/2510.07226v1) (tool)
- [umiacs.umd.edu](https://www.umiacs.umd.edu/news-events/news/report-ai-use-newspapers-widespread-uneven-and-rarely-disclosed) (tool)
- [arxiv.org](https://arxiv.org/pdf/2510.18774) (tool)
- [arxiv.org](https://arxiv.org/abs/2406.07016) (tool)
- [github.com](https://github.com/berenslab/llm-excess-vocab) (tool)
- [publications.ut-capitole.fr](https://publications.ut-capitole.fr/id/eprint/52414/1/wp_tse_1722.pdf) (tool)
- [arxiv.org](https://arxiv.org/abs/2608.10715) (tool)
- [arxiv.org](https://arxiv.org/html/2608.10715v1) (tool)
- [arxiv.org](https://arxiv.org/abs/2403.07183) (tool)
- [AI Now Writes as Many Online Articles as Humans — Five Percent](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do) (openai)
- [ahrefs.com](https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated) (tool)
- [originality.ai](https://originality.ai/blog/dead-internet-tracker-july) (tool)
- [arxiv.org](https://arxiv.org/html/2412.18148v1) (tool)
- [arxiv.org](https://arxiv.org/abs/2504.08755) (tool)
- [arxiv.org](https://arxiv.org/abs/2510.03154) (tool)
- [bfi.uchicago.edu](https://bfi.uchicago.edu/working-papers/artificial-writing-and-automated-detection?trk=article-ssr-frontend-pulse_little-text-block) (tool)
- [pangram.com](https://www.pangram.com/blog/humanizers-aug-25) (tool)
- [pangram.com](https://www.pangram.com/research/papers) (tool)
- [arxiv.org](https://arxiv.org/abs/2607.27183) (tool)
- [arxiv.org](https://arxiv.org/abs/2405.07940) (tool)
- [sciencedirect.com](https://www.sciencedirect.com/science/article/pii/S2667305326000840) (tool)
- [gptzero.me](https://gptzero.me/news/how-does-gptzero-compare-to-pangram-v3-2-latest-benchmark-results) (tool)
- [pewresearch.org](https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact) (tool)
- [nber.org](https://www.nber.org/papers/w32966) (tool)
- [nber.org](https://www.nber.org/papers/w35677) (tool)
- [https://arxiv.org/pdf/2604.26965](https://arxiv.org/pdf/2604.26965) (openai)
- Epoch (mcp)
  > Poll 'mar_2026' (2026-03): 59 rows.
- [epoch.ai](https://epoch.ai/data/polling_on_ai_usage_mar_2026.csv) (tool)
- [Link Rot and Digital Decay on Government, News and Other Webpages | Pew Research Center](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears) (openai)

## Question Details

This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. (pewresearch.org) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.

### Resolution Criteria

Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. (pewresearch.org) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. (pewresearch.org) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.

### Fine Print

“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. (pewresearch.org) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.
