Back to question

Forecast report

What percent of webpages will be significantly written or edited by AI in 2030?

GeneratedOctober 11, 2026 at 11:36 AM UTC
ResolutionNot specified
Question typeNumeric
Sources50

Forecast

Median forecast: 29; 80% interval: 15 to 55.

Distribution

0.0%0.6%1.2%1.8%2.4%025507510080% intervalMedian

Analysis

TL;DR

My median forecast is 29% of English-language webpages, with a predictive mean of 31.7% (forecast calculation). The central 80% predictive interval is 15%–54.8%, and the probability of exceeding 50% is 13% (forecast calculation). This forecast includes old pages and is conditional on a qualifying numerical resolution, not annulment (resolution rules).

Context

<a id="resolution-rules"></a>The information cutoff is October 11, 2026, at 11:20 UTC. The supplied rules target all qualifying English-language pages in the latest qualifying snapshot from calendar year 2030, including older and undated pages; historical observations inform the forecast but cannot resolve it early. If no comparable study appears by December 31, 2032, the question annuls rather than resolving to zero (supplied resolution rules).

The outcome is a study’s page-level classification rate, not the percentage of words generated by AI or proof of individual authorship. Pew’s baseline measures body text and explicitly treats its detector as probabilistic (Pew methodology).

Evidence

The historical backbone is Pew’s study published August 20, 2026: 10,000 English-language pages from each of 49 Common Crawl snapshots, totaling 490,000 page observations. It flags meaningful AI writing or editing at an Open Pangram/EditLens score of at least 0.2. Only 10%–15% of pages have identifiable publication dates, and Pew says that subset is nonrandom (Pew methodology). (pewresearch.org)

The complete all-page history follows. Values are percentages of sampled pages, extracted from the article’s embedded chart data—not reconstructed from domain averages. Coverage runs from January 27, 2021, through July 18, 2026; the retrieved article metadata marks September 17, 2026, as its modification date (Pew all-page data). (pewresearch.org)

Snapshot dateAI-positive share (%)Snapshot dateAI-positive share (%)
2021-01-271.182021-03-031.13
2021-04-201.102021-05-171.12
2021-06-231.082021-08-011.11
2021-09-191.072021-10-201.13
2021-11-301.092022-01-241.10
2022-05-251.042022-07-061.04
2022-08-161.022022-09-291.07
2022-12-081.152023-02-041.35
2023-03-201.532023-06-091.89
2023-09-212.302023-12-102.78
2024-03-013.222024-04-133.55
2024-05-183.492024-06-243.49
2024-07-243.402024-08-043.80
2024-09-203.832024-10-124.34
2024-11-024.512024-12-104.92
2025-01-224.862025-02-064.92
2025-03-234.982025-04-225.07
2025-05-124.932025-06-224.86
2025-07-144.782025-08-135.24
2025-09-125.492025-10-156.01
2025-11-156.292025-12-116.64
2026-01-256.922026-02-197.28
2026-03-158.092026-04-208.58
2026-05-198.862026-06-068.98
2026-07-189.60——

Pew warns that the early readings include false positives. Its comparison of 62,370 pages from seven crawls found 96% agreement with commercial Pangram 3.3, but agreement between detectors is not ground-truth accuracy (Pew methodology). I therefore exclude pre-ChatGPT observations from growth fitting and forecast the comparable reported classification rate, without an exact correction for latent authorship.

My equal-weighted, calendar-date fits show acceleration. The linear slope is 2.1 percentage points per year across all post-ChatGPT observations, versus 4.8 over the latest year. Fixed-ceiling sigmoid fits produce very different futures despite similar historical errors. These are my calculations from the full series, not Pew predictions (underlying observations).

Extrapolation to December 31, 2030Snapshot observationsProjected share (%)
Linear, December 2022 onward3517.5
Linear, January 2025 onward1923.9
Linear, July 2025 onward1330.9
Sigmoid, assumed eventual ceiling 25%3522.4
Sigmoid, assumed eventual ceiling 40%3531.2
Sigmoid, assumed eventual ceiling 60%3539.5
Sigmoid, assumed eventual ceiling 80%3545.4
Sigmoid, assumed eventual ceiling 100%3549.9

The ceiling is not identified by the history. I use these fits as checks on the reasonable range, not as models whose small residuals justify narrow forecast intervals.

New-content evidence supports a higher incoming AI share, but also a slowdown in adoption. Graphite’s May 15, 2026 study sampled 55,400 English articles and listicles published from January 2020 through March 2026, requiring article metadata and at least 100 words. It averaged three detectors’ primarily-AI classifications; its full quarterly history appears below (study and definitions, complete source data). (graphite.io)

Publication yearQ1 (%)Q2 (%)Q3 (%)Q4 (%)
20200.970.900.950.64
20211.041.230.891.18
20221.542.373.854.60
202314.0225.1433.6935.92
202438.2337.2241.8147.04
202549.6144.5849.2250.90
202649.94———

This is a plateau in a selected article population, not the entire web. Its primarily-AI definition also differs from meaningful editing, and the study did not validate detection after substantial human rewriting (Graphite limitations). An April 14, 2026 preprint provides weaker corroboration: its mid-2025 estimate was 35% AI-generated or assisted, but it used first archival dates, a one-URL-per-host cap, and long paragraphs; the retained analytical sample size is not clearly reported (Dolezal et al.).

Old material slows the transition. Pew’s May 17, 2024 digital-decay study sampled approximately one million URLs from 2013–2023 and checked accessibility during October 12–November 6, 2023. It found 38% of the 2013 cohort inaccessible a decade later. That measures disappearance, not unchanged text, and is not an English-page renewal estimate (findings, methodology). (pewresearch.org)

Crawler discovery is also not publication. Common Crawl’s September 19, 2026 release contained 2.17 billion captures collected September 4–17, including 587.2 million previously unvisited URLs. Those multilingual captures cannot be annualized into an English-text replacement rate (Common Crawl release).

Existing-page rewriting is a real channel. Shopify reported on March 24, 2026 that merchants created 19.8 million product descriptions in the first three weeks after its Winter ’26 release, which launched December 10, 2025. Merchant sample size, unique published URLs, language, and old-versus-new descriptions were not disclosed (Shopify production figures, release date). WordPress.com’s March 20 announcement, updated May 15, enables agents to update published pages, but requires enabled permissions and approval (WordPress.com announcement). These sources support scalable rewriting, not blanket conversion of the legacy web.

Economic incentives restrain a simple flood scenario. Google’s guidance, updated October 1, 2026, targets mass generation without user value, not useful AI-assisted writing generally. Search exclusion also does not mean removal from a public crawl. I treat this as a brake on indiscriminate publishing, not a cap on AI prevalence (Google guidance).

Detector evolution moves the result in both directions. Pangram’s July 29, 2026 vendor report scored the same 14,990 substantially edited student texts with two detector versions: Mixed-or-AI classifications rose from 3,113 to 8,789, or 21% to 58.6%. This is a selected benchmark, not a web audit or a comparison at Pew’s threshold (Pangram 4 report). (arxiv.org) Conversely, a September 25 experiment on 400 deliberately constructed responses reduced EditLens detection from 55% to 6% through base-model orchestration, using a separately calibrated threshold. That establishes an evasion mechanism, not its prevalence online (Dhingra and Pruthi).

<a id="forecast-model"></a>I combine this evidence through an effective stock-adjustment model:

dsdt=λ [q(t)−s(t)].\frac{ds}{dt}=\lambda\,[q(t)-s(t)].

Here, ss is the all-page AI-positive fraction, qq is the comparable classified fraction of entering or substantially rewritten material, and λ\lambda is the annual adjustment rate. It summarizes additions, removals, rewriting, and crawl composition—not a measured page-death hazard. I start at the measured baseline and let qq rise linearly within each scenario. The following parameters, weights, and uncertainty spreads are my judgments, encoded in the accompanying calculation (forecast model).

ScenarioWeightAdjustment/yearFlow share, initial → finalTiming-averaged mean (%)Within-scenario SD, pp
Downturn or adverse detection/coverage shift3%——8.04.5
Slow renewal and persistent plateau25%0.0835% → 45%18.75.5
Mainstream hybrid diffusion52%0.1636% → 65%30.37.5
Rapid rewriting and publishing17%0.3045% → 85%52.111.0
Exceptional AI-heavy composition3%0.6050% → 97%76.110.0

The central scenario begins with growth of 4.2 percentage points per year, consistent with recent acceleration without extending exponential growth indefinitely. The slow scenario preserves the long-window extrapolation and persistent human stock. The upper scenarios preserve materially different futures involving faster rewriting, automated publication, and stronger editing detection; the downturn scenario allows a future measurement below today’s baseline (forecast model).

I integrate snapshot timing over March 31, July 18, and December 31 of 2030, with weights of 15%, 35%, and 50%. These are assumptions, not announced crawls. Each scenario has bounded beta uncertainty, followed by a nominal 10,000-page binomial sample. Reporting precision receives weights of 60% for hundredths, 10% for tenths, and 30% for whole percentages; this puts extra mass into buckets ending at integers when only rounded results are available (forecast model).

The resulting median is 29%, the mean 31.7%, and the central 90% interval 12%–64.3%. It assigns 13% probability above 50%, 4% at or below 10%, and 1% above 80%. These are predictive uncertainties about the eventual reported result, not sampling confidence intervals. No prediction-market signal was used (forecast calculation).

What's non-obvious

The physical web and the sampled web do not age at the same rate. Common Crawl says it recrawls slowly changing pages less often so it can focus on newly discovered and frequently changing pages. A persistent human-written page can therefore remain online while becoming less represented in snapshots (coverage explanation). Selection also explains some apparently conflicting AI estimates: a September 30 preprint’s audit of 10,000 raw captures found AI-labeled documents survived FineWeb filtering 2.3 times as often as human-labeled documents. Filtered-token estimates cannot be converted into all-page estimates with a fixed multiplier (filtering audit).

Human ideas can still produce AI-written pages. An October 5 paper separates idea provenance from prose provenance and demonstrates that detectors aimed at these concepts behave differently (IdeaLens study). This question counts substantial AI wording or editing even when a human supplied the argument. A future detector that measures only AI-origin ideas would change the target rather than improve its measurement. Likewise, old publication dates do not protect pages from becoming AI-positive through later rewriting (Pew’s editing definition).

Uncertainties

The biggest gaps are linked rather than independent:

  • The rate of substantial rewriting on previously published English URLs is unmeasured. Platform generation counts do not establish acceptance, publication, unique-page counts, or human-to-AI conversion (Shopify disclosure).
  • Detector upgrades need a representative, paired web evaluation with recorded writing histories. Current editing and evasion benchmarks do not supply that bridge, and detector agreement does not establish accuracy (Pangram report, Pew validation).
  • Crawl coverage and page-type composition can change independently of authorship. Login-only and paywalled content is underrepresented, while changing pages receive more crawl attention (Pew coverage limits, Common Crawl coverage).
  • Study availability, snapshot timing, and methodological comparability remain resolution risks. An unavailable or nonqualifying estimate leads to annulment, not a numerical outcome at the bottom of the distribution (resolution rules).

A repeated-URL panel combining publishing logs with old and new detectors would close the largest gap. Until then, the evidence supports a rising all-page share, but it does not identify a single renewal rate or detector correction. The forecast retains substantial probability on both slow accumulation and rapid rewriting instead of hiding that uncertainty inside a precise point estimate (forecast model).

Sources

  1. Domain Expert Search · mcp

    Found 7 domain experts for 'AI-generated web text measurement AI detector accuracy substantial AI editing Pangram and Common Crawl sampling biases':

  2. News · mcp

    Found 12 merged articles (asknews: 6, perigon: 6, both: 0).

  3. business-standard.com · tool
  4. Pew says 10% of the web now shows signs of AI authorship · anthropic
  5. medium.com · tool
  6. How Much of the Internet Is Written With AI? | Pew Research Center · openai
  7. socialmediaexaminer.com · tool
  8. nypost.com · tool
  9. flowingdata.com · tool
  10. pewresearch.org · tool
  11. thedailystar.net · tool
  12. theverge.com · tool
  13. platform.theverge · tool
  14. timesnownews.com · tool
  15. sfgate.com · tool
  16. varonis.com · tool
  17. theblaze.com · tool
  18. thedailytechfeed.com · tool
  19. nymag.com · tool
  20. news18.com · tool
  21. finance.yahoo.com · tool
  22. Claude Code · e2b

    Job coding_whiz_job_da4d59a737 done after 215173ms.

  23. pewresearch.org · tool
  24. pewresearch.org · tool
  25. Methodology | Pew Research Center · openai
  26. Domain Expert Research Task · mcp

    Job domain_expert_research_task_a0696a4375 done after 426611ms.

  27. arxiv.org · tool
  28. arxiv.org · tool
  29. arxiv.org · tool
  30. arxiv.org · tool
  31. AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors | Epoch AI · openai
  32. arxiv.org · tool
  33. arxiv.org · tool
  34. arxiv.org · tool
  35. github.com · tool
  36. arxiv.org · tool
  37. shopify.com · tool
  38. abcailab.com · tool
  39. changelog.shopify.com · tool
  40. shopify.com · tool
  41. shopify.com · tool
  42. help.shopify.com · tool
  43. wordpress.com · tool
  44. wordpress.com · tool
  45. wordpress.com · tool
  46. wordpress.com · tool
  47. siteground.com · tool
  48. wordpress.org · tool
  49. The Wix State of Websites Report in 2026 · openai
  50. support.wix.com · tool

Question Details

Description

This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.

Resolution Criteria

Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.

Fine Print

“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.