# What percent of webpages will be significantly written or edited by AI in 2030?

Canonical URL: https://preseen.com/reports/0e135a67-fddd-4061-b8e4-45479617dafa/what-percent-of-webpages-will-be-significantly-written-or-edited-by-ai-in-2030
Markdown URL: https://preseen.com/reports/0e135a67-fddd-4061-b8e4-45479617dafa/markdown

## Forecast

Median forecast: 29.5; 80% interval: 13 to 58.

Generated: October 9, 2026 at 11:41 AM UTC
Forecast model: gpt-6.1-sol
Research model: gpt-5.6-sol+gpt-6.1-sol+claude-opus-5-5

## Analysis

## TL;DR
My median forecast is **29%** of sampled English-language webpages in the qualifying 2030 snapshot ([forecast calculation](#evidence)). The central 80% predictive interval is 13%–57.7%, with a 15% probability of a reported share above 50% ([distribution calculation](#evidence)). This forecasts the entire page stock, conditional on a qualifying study, not the share of newly published articles ([target and baseline](#context)).

## Context
The evidence cutoff is October 9, 2026, at 11:25 UTC. The latest verified Pew all-page observation is 9.60% for July 18, 2026, rounded to 10% in its August 20 report ([Pew study](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)). The target is the comparable reported classification share in a snapshot from calendar year 2030. Historical observations inform this forecast; they do not resolve it.

Pew sampled 10,000 English-language pages from each of 49 crawls spanning January 2021–July 2026: 490,000 page observations. It analyzed body text with Open Pangram and classified scores of at least 0.2 as meaningful AI authorship or editing. Only about 10%–15% of pages had detectable publication dates, and that subset was nonrandom ([methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)). The numerical distribution is conditional on a qualifying result; annulment is not a result of zero.

## Evidence
The historical backbone is the complete all-page series below, recovered from the embedded chart in the report. These are crawl dates, not publication dates; every row has the same sample size stated above. The vintage is the report available at the evidence cutoff ([Pew chart data](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)).

| Crawl date | Classified share |
|---|---:|
| 2021-01-27 | 1.18% |
| 2021-03-03 | 1.13% |
| 2021-04-20 | 1.10% |
| 2021-05-17 | 1.12% |
| 2021-06-23 | 1.08% |
| 2021-08-01 | 1.11% |
| 2021-09-19 | 1.07% |
| 2021-10-20 | 1.13% |
| 2021-11-30 | 1.09% |
| 2022-01-24 | 1.10% |
| 2022-05-25 | 1.04% |
| 2022-07-06 | 1.04% |
| 2022-08-16 | 1.02% |
| 2022-09-29 | 1.07% |
| 2022-12-08 | 1.15% |
| 2023-02-04 | 1.35% |
| 2023-03-20 | 1.53% |
| 2023-06-09 | 1.89% |
| 2023-09-21 | 2.30% |
| 2023-12-10 | 2.78% |
| 2024-03-01 | 3.22% |
| 2024-04-13 | 3.55% |
| 2024-05-18 | 3.49% |
| 2024-06-24 | 3.49% |
| 2024-07-24 | 3.40% |
| 2024-08-04 | 3.80% |
| 2024-09-20 | 3.83% |
| 2024-10-12 | 4.34% |
| 2024-11-02 | 4.51% |
| 2024-12-10 | 4.92% |
| 2025-01-22 | 4.86% |
| 2025-02-06 | 4.92% |
| 2025-03-23 | 4.98% |
| 2025-04-22 | 5.07% |
| 2025-05-12 | 4.93% |
| 2025-06-22 | 4.86% |
| 2025-07-14 | 4.78% |
| 2025-08-13 | 5.24% |
| 2025-09-12 | 5.49% |
| 2025-10-15 | 6.01% |
| 2025-11-15 | 6.29% |
| 2025-12-11 | 6.64% |
| 2026-01-25 | 6.92% |
| 2026-02-19 | 7.28% |
| 2026-03-15 | 8.09% |
| 2026-04-20 | 8.58% |
| 2026-05-19 | 8.86% |
| 2026-06-06 | 8.98% |
| 2026-07-18 | 9.60% |

The early positives are not proof of widespread pre-ChatGPT authorship: Pew identifies likely false positives in those samples ([methodology](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

**The recent stock trend accelerates rather than plateaus.** Anchored linear extrapolations over a further 4.2 years range from 18% using all 35 post-launch observations to 32.9% using the latest seven. These are my calculations from the full series, not Pew forecasts ([underlying observations](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)). Window sensitivity is more useful here than a narrow regression confidence interval.

Older material provides a brake. Pew’s May 17, 2024 survival study sampled just under one million pages from annual 2013–2023 archives and checked accessibility in October 2023. It found that 38% of the 2013 cohort was inaccessible a decade later ([digital-decay study](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/)). A July 15, 2026 Common Crawl paper, using 51 crawls over 2020–2025, found that a persistent core plus a changing outer population fit better than homogeneous turnover; its estimated core was about 40% at domain level ([crawl-persistence paper](https://arxiv.org/html/2607.13636v1)). Neither study measures how often surviving page bodies become AI-edited. I therefore use them to justify heterogeneous scenarios, not to manufacture a measured text-replacement rate.

Incoming content is much more AI-heavy than the stock. Dolezal and colleagues’ April 14, 2026 preprint estimated about 35% AI-generated or AI-assisted first-archived websites by mid-2025. Its sampling covered August 2022–May 2025, targeted 10,000 URLs per monthly interval, restricted hosts, and analyzed an eligible longest paragraph; I could not verify the final retained English-language sample size ([study and sampling appendix](https://arxiv.org/html/2604.26965v1)). Graphite’s 2026 analysis sampled 55,400 articles and listicles dated January 2020–March 2026, with detector evaluations from April–May 2026. Its latest quarter had 49.9% primarily AI-generated articles ([article study](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do)). These are selected publishing populations, not interchangeable all-page estimates.

The strongest upward signal is easier editing of existing sites. WordPress.com introduced an assistant that can rewrite and translate page text, then extended opt-in access to all current paid plans in May 2026 ([assistant capabilities](https://wordpress.com/blog/2026/02/17/wordpress-ai-assistant/), [availability update](https://wordpress.com/blog/2026/05/08/changelog-assistant-studio-code/)). Availability is not measured adoption. But it creates a conversion channel that does not require old URLs to disappear. The opposing incentive is search-quality enforcement: Google’s guidance, updated October 1, 2026, warns against mass generation without user value, while allowing useful AI assistance ([Google guidance](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content)). I read this as a brake on low-value publishing, not a ban on AI editing or a guarantee that rejected pages leave the crawl.

Measurement can move the result substantially. In the July 29, 2026 Pangram 4 report, a controlled evaluation of 14,990 substantially AI-edited texts produced positive classifications—Mixed plus AI—on 21% with Pangram 3.3.2 and 58.6% with Pangram 4. This is a vendor-led editing benchmark, not a representative webpage audit or a direct test of Pew’s open checkpoint ([Table 5](https://arxiv.org/html/2607.27183v1)). It supports a meaningful measurement tail, not multiplying the baseline by the benchmark’s recall ratio.

A September 30, 2026 preprint reported 31.1% AI-labeled tokens in an August sample of 5,000 FineWeb-filtered documents. Its separate filtering audit followed 10,000 raw documents and found that the pipeline retained AI-labeled documents 2.3 times as often as human-labeled documents ([methods and filtering audit](https://arxiv.org/html/2609.40295v1)). Token weighting and AI-favoring filtration explain why this higher estimate cannot replace the page-count baseline. Nor can document-retention ratios be mechanically applied to token shares.

I turn this evidence into a stock-flow model:

$$
\frac{ds}{dt}=\lambda\bigl(q(t)-s(t)\bigr).
$$

Here s is the classified share of the whole snapshot, q is the classified share of entering or substantially rewritten content, and λ is an effective annual composition-refresh rate. It includes creation, rewriting, disappearance and crawl recomposition. It is not a URL-death rate. I start at the observed 0.096 and use a 4.2-year horizon, representing a relatively late qualifying snapshot. Earlier collection dates are covered by the uncertainty distributions.

The following are my forecasting assumptions, not fitted empirical rates. Within each growth scenario, q increases linearly between the listed endpoints. The code solves the equation without rounding its terminal value.

| Scenario | Weight | Annual refresh rate λ | Incoming positive share, start → end | Calculated center | Beta concentration |
|---|---:|---:|---:|---:|---:|
| Plateau or measurement downside | 10% | Stress case | Stress case | 12% | 25 |
| Restrained growth | 20% | 9% | 35% → 45% | 19% | 30 |
| Ordinary diffusion | 50% | 15% | 45% → 65% | 31.3% | 28 |
| Rapid publishing and editing | 15% | 30% | 50% → 85% | 53.6% | 20 |
| Very rapid automation | 5% | 45% | 60% → 95% | 72% | 16 |

The ordinary path begins with growth of 5.3 percentage points per year, consistent with the recent observed pace. Holding its incoming share constant instead of increasing it lowers its terminal center to 26.1%. Thus continued adoption matters; turnover alone does not produce the central result ([model calculation](#evidence)). The rapid paths require a genuine change beyond the observed pace, so they receive minority weight.

Each scenario gets a beta distribution on the permitted range. For a fractional mean m and concentration k, its parameters are mk and (1−m)k. Concentrations are uncertainty settings, not sample sizes; they give within-scenario standard deviations of roughly 6–11 percentage points. The downside component is a directly specified stress distribution. These widths retain disagreement across trend windows, publishing populations and measurement methods rather than treating overlapping archives and detector families as independent confirmation.

Finally, I assign 80% weight to sufficiently granular reporting and 20% to whole-percentage reporting. This is a reporting assumption, not an observed future practice. Whole-number results go into the buckets ending at those integers. The resulting mean is 32.3%, the median is 29%, and the central 90% interval is 9%–67.2% ([distribution calculation](#evidence)).

## What's non-obvious
**A plateau in new-content AI prevalence would not be a plateau in the whole-web share.** AI-heavy material keeps accumulating while older material remains in the denominator. The reverse mistake is treating persistent URLs as permanently human-written: substantial AI revisions can change their classification without changing their address. Neither new-article percentages nor link survival alone answers this question ([article population](https://graphite.io/five-percent/research/ai-now-writes-as-many-online-articles-as-humans-do), [persistence evidence](https://www.pewresearch.org/data-labs/2024/05/17/when-online-content-disappears/), [editing capability](https://wordpress.com/blog/2026/02/17/wordpress-ai-assistant/)).

The measured percentage can also rise without an equally large rise in actual AI use. Better mixed-authorship detection can recover missed editing, while corpus filters can make a selected dataset look AI-heavy. Those are different mechanisms, and neither supplies a universal correction factor ([editing benchmark](https://arxiv.org/html/2607.27183v1), [filtering audit](https://arxiv.org/html/2609.40295v1)).

## Uncertainties
The largest missing evidence is a representative panel that separates genuinely new pages, substantive revisions at existing URLs, and changes in crawl inclusion. The stock-flow parameters remain weakly identified without it. The useful additional data would be:

- Repeated page-body observations with reliable creation dates and documented substantial revisions. URL persistence and first archival do not supply this information ([crawl-persistence analysis](https://arxiv.org/html/2607.13636v1)).
- Independent, representative human and mixed-authorship validation sets linking future detectors to the same substantial-editing concept. Controlled editing prompts do not establish broad-web sensitivity or false-positive rates ([Pew’s measurement caveats](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/), [Pangram evaluation](https://arxiv.org/html/2607.27183v1)).
- Stable information about future crawl priorities, exclusions, extraction and snapshot timing. These can move the reported share even when author behavior changes smoothly ([crawl-selection analysis](https://arxiv.org/html/2607.13636v1)).

Ordinary sampling error is much smaller than these gaps: an independent sample of 10,000 pages near the ordinary scenario center would have a standard error of about 0.46 percentage points. That calculation excludes clustering and systematic measurement error ([sample-size basis](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)).

The current baseline is not a hard floor. Recalibration and changing snapshot composition can produce a lower qualifying result; the distribution retains a 6% probability below 10% ([distribution calculation](#evidence)). Under the client’s resolution criteria, failure to obtain a sufficiently comparable study of a 2030 snapshot by December 31, 2032 annuls the question. The array is normalized conditional on numerical resolution and assigns no annulment probability to zero.

## Sources

- Domain Expert Search (mcp)
  > Found 7 domain experts for 'AI generated web content detection substantial human AI editing detector reliability Pangram open model representative Common Crawl measurement':
- AskNews (mcp)
  > Found 8 articles:
- Claude Code (e2b)
  > Job coding_whiz_job_aec2ca62da done after 291956ms.
- [How Much of the Internet Is Written With AI? | Pew Research Center](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai) (openai)
- [Methodology | Pew Research Center](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content) (openai)
- Domain Expert Research Task (mcp)
  > Job domain_expert_research_task_2c0c6ad2b3 done after 409151ms.
- [\[2607.27183\] Pangram 4 Technical Report](https://arxiv.org/abs/2607.27183) (openai)
- [How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text](https://arxiv.org/abs/2609.40295) (openai)
- Epoch (mcp)
  > Poll 'mar_2026' (2026-03): 59 rows.
- [epoch.ai](https://epoch.ai/data/polling_on_ai_usage_mar_2026.csv) (tool)
- Anthropic Economic Index Data (mcp)
  > AUTOMATION VS AUGMENTATION - share of CLASSIFIED Claude conversations
- Cloudflare Radar (mcp)
  > HTTP traffic share by bot_class (%)
- arXiv (mcp)
  > arXiv ID: 2604.26965
- [The Impact of AI-Generated Text on the Internet](https://ai-on-the-internet.github.io/) (openai)
- [huggingface.co](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) (tool)
- [huggingface.co](https://huggingface.co/cardiffnlp/twitter-roberta-base-sentiment-latest) (tool)
- [arxiv.org](https://arxiv.org/pdf/2605.00087) (tool)
- [arxiv.org](https://arxiv.org/pdf/2605.00087v1) (tool)
- xai (mcp)
  > Sources:
- [aidetectors.io](https://www.aidetectors.io/blog/do-ai-humanizers-work) (tool)
- [franceinfo.fr](https://www.franceinfo.fr/replay-radio/le-vrai-du-faux/la-desinformation-russe-explose-avec-850-de-faux-reportages-en-plus-en-un-an-selon-newsguard_8194607.html) (tool)
- [san.com](https://san.com/cc/inside-a-newspaper-chains-transformation-from-local-news-to-ai-generated-content) (tool)
- [mobileappdaily.com](https://www.mobileappdaily.com/products/best-ai-website-builders) (tool)
- [tenten.co](https://tenten.co/shopify/ai-native-cms-crm-pim-2026) (tool)
- [easync.io](https://easync.io/articles/best-ai-dropshipping-website-builder) (tool)
- [searchenginejournal.com](https://www.searchenginejournal.com/googles-spam-update-now-reaches-ai-answers-enforcement-is-hard/580535) (tool)
- [keywordseverywhere.com](https://keywordseverywhere.com/news/google-algorithm-updates/ai-slop) (tool)
- [blogs.bing.com](https://blogs.bing.com/search/2026/5/Keeping-Trusted-Content-Visible-in-an-AI-Powered-Search-World) (tool)
- [therepository.email](https://www.therepository.email/wordpress-mcp-adapter-now-available-as-a-canonical-plugin-on-wordpress-org) (tool)
- [twipemobile.com](https://www.twipemobile.com/kleine-zeitungs-plan-for-a-web-run-by-ai-agents) (tool)
- [auto-post.io](https://auto-post.io/blog/automate-blog-ops-with-ai-agents) (tool)
- [originality.ai](https://originality.ai/blog/publisher-embraces-ai-case-study) (tool)
- [newsguardtech.com](https://www.newsguardtech.com/special-reports/ai-tracking-center) (tool)
- [learn.microsoft.com](https://learn.microsoft.com/en-us/microsoft-365/copilot/manage-public-web-access) (tool)
- [marketing-interactive.com](https://www.marketing-interactive.com/wordpress-lets-ai-agents-write-design-and-manage-your-site) (tool)
- [searcheye.io](https://searcheye.io/blog/best-ai-website-builders) (tool)
- [mobileappdaily.com](https://www.mobileappdaily.com/products/best-ai-seo-tools) (tool)
- [make.wordpress.org](https://make.wordpress.org/ai/2026/08/18/whats-new-in-ai-1-3-0) (tool)
- [newsguardtech.com](https://www.newsguardtech.com/press/newsguard-launches-real-time-ai-content-farm-detection-datastream-to-counter-onslaught-of-ai-slop-in-news) (tool)
- [humanizereval.com](https://humanizereval.com/) (tool)
- [thenicheguru.com](https://thenicheguru.com/tools/best-ai-website-builder) (tool)
- [xseek.io](https://www.xseek.io/blogs/articles/best-seo-tools-for-ai-optimized-content-in-2026) (tool)
- [disinfocode.eu](https://www.disinfocode.eu/reports/microsoft-bing/9?commitmentId=407&chapterId=84) (tool)
- [humanizerbench.com](https://humanizerbench.com/blog/october-2026-ai-humanizer-rankings) (tool)
- [itechguides.com](https://www.itechguides.com/best/ai-content-creation-and-optimization-software) (tool)
- [blockchain.news](https://blockchain.news/news/ai-website-builders-ecommerce-2026-comparison) (tool)
- [wordpress.org](https://wordpress.org/plugins/ai) (tool)
- [elsop.com](https://www.elsop.com/google-spam-policy-ai-answers) (tool)
- [umatechnology.org](https://umatechnology.org/best/ai-website-builders) (tool)
- [lafactory.com](https://lafactory.com/wordpress-ai-abilities-api-mcp) (tool)

## Question Details

This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. (pewresearch.org) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.

### Resolution Criteria

Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. (pewresearch.org) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. (pewresearch.org) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.

### Fine Print

“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. (pewresearch.org) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.
