Back to question

Forecast report

What percent of webpages will be significantly written or edited by AI in 2030?

GeneratedOctober 5, 2026 at 11:28 AM UTC
ResolutionNot specified
Question typeNumeric
Sources50

Forecast

Median forecast: 30; 80% interval: 14.5 to 56.5.

Distribution

0.0%0.5%1.0%1.5%2.0%025507510080% intervalMedian

Analysis

TL;DR

My central forecast is 30%, with an expected value of 32.8% (calculation below). The central 80% predictive range is 14%–56%, and the probability of a result above 50% is 15% (distribution below). This forecast counts all qualifying webpages, not just new content, and is conditional on a numerical resolution rather than annulment (interpretation below).

Context

The target is a page-count share in the latest qualifying calendar-2030 snapshot. Old pages remain in the denominator; tokens, traffic, newly published articles, and non-textual AI use do not replace that denominator. I forecast the study’s classification statistic, not proven authorship, following the concept defined by Pew’s baseline and methodology.

The directly comparable starting point is 9.60% in the July 2026 all-page sample, rounded to 10% in the headline. The separate 35% estimate concerns dated post-ChatGPT pages, a nonrandom subset, and cannot serve as the all-web baseline (Pew, August 20, 2026; sampling caveat). (pewresearch.org)

Evidence

The historical backbone is Pew’s random sample of 10,000 English-language pages from each of 49 crawls spanning January 2021–July 2026: 490,000 page observations, not necessarily distinct pages. It classified extracted body text using Open Pangram at a score of at least 0.2. Only 10%–15% of pages exposed publication dates (Pew methodology, August 20, 2026). (pewresearch.org)

The complete all-page history follows. Values are percentages of sampled pages classified positive, decoded from the published chart rather than reconstructed from selected domain averages. The source was published August 20, 2026; its top-level modification metadata gives September 17, 2026 (Pew chart and article).

Crawl datePositive %Crawl datePositive %Crawl datePositive %
2021-01-271.182021-03-031.132021-04-201.10
2021-05-171.122021-06-231.082021-08-011.11
2021-09-191.072021-10-201.132021-11-301.09
2022-01-241.102022-05-251.042022-07-061.04
2022-08-161.022022-09-291.072022-12-081.15
2023-02-041.352023-03-201.532023-06-091.89
2023-09-212.302023-12-102.782024-03-013.22
2024-04-133.552024-05-183.492024-06-243.49
2024-07-243.402024-08-043.802024-09-203.83
2024-10-124.342024-11-024.512024-12-104.92
2025-01-224.862025-02-064.922025-03-234.98
2025-04-225.072025-05-124.932025-06-224.86
2025-07-144.782025-08-135.242025-09-125.49
2025-10-156.012025-11-156.292025-12-116.64
2026-01-256.922026-02-197.282026-03-158.09
2026-04-208.582026-05-198.862026-06-068.98
2026-07-189.60————

These observations support acceleration, but not one uniquely identified growth curve. I fitted straight lines anchored at the latest observation and bounded log-odds trends with assumed ceilings. The following are my December-2030 calculations from the full Pew series, not published predictions.

Fitting windowObservationsAnchored linearLog-odds, 50% ceilingLog-odds, 100% ceiling
Since February 20233422.35%37.77%54.13%
Last three years3123.26%37.21%52.32%
Last two years2524.90%38.25%54.55%
Last year1231.34%43.56%69.56%

The ceilings are assumptions. These fits reuse the same evidence, so they are sensitivity checks, not independent forecasts to average. Their disagreement is much larger than the approximately 0.3-percentage-point sampling standard error calculated for the latest sample under an independent-binomial approximation (sample design).

Evidence about incoming content supplies a second anchor. Graphite’s commercially produced study selected 55,400 dated English articles and listicles, at least 100 words long, published January 2020–March 2026. It averaged three detector classifications. Its complete quarterly history shows rapid adoption followed by a plateau near half, rather than uninterrupted exponential growth (Graphite, May 15, 2026; March crawl and April–May detector vintage). (graphite.io)

Publication yearQ1 positive %Q2 positive %Q3 positive %Q4 positive %
20200.970.900.950.64
20211.041.230.891.18
20221.542.373.854.60
202314.0225.1433.6935.92
202438.2337.2241.8147.04
202549.6144.5849.2250.90
202649.94———

That study targets primarily AI-generated articles, not Pew’s broader substantial-editing concept or all-page denominator. I use its plateau to support a slowdown scenario, not a universal ceiling (Graphite’s classifications and limitations).

Legacy content also matters. Pew sampled just under one million URLs across annual 2013–2023 crawls and checked accessibility in October 2023. It found 38% of the 2013 cohort inaccessible, leaving most accessible a decade later. That is evidence of persistence, but it measures URL accessibility—not unchanged text or denominator weight after new pages enter (Pew, May 17, 2024). (pewresearch.org)

Publishing tools provide a route around that legacy constraint. WordPress.com opened opt-in AI editing assistance across its current paid plans in May 2026. This is feature availability on WordPress.com, not measured uptake across the entire WordPress ecosystem (official May 8 announcement). Search incentives pull the other way: Google warns against scaled page generation without user value, rather than banning useful AI assistance. I read this as a headwind to commodity publishing, not a numerical cap on crawl prevalence (Google guidance). (wordpress.com)

Measurement changes deserve substantial weight. In a vendor-controlled evaluation of 14,990 substantially AI-edited student essays, the share labeled Mixed or AI rose from 20.8% under Pangram 3.3.2 to 58.6% under Pangram 4. This demonstrates that improved mixed-authorship sensitivity can produce large reclassification. It does not supply a correction to a random webpage sample or Pew’s open-model threshold (Pangram technical report, July 29, 2026). (arxiv.org)

I translate these forces into an effective stock-renewal model:

dpdt=λ [q(t)−p(t)].\frac{dp}{dt}=\lambda\,[q(t)-p(t)].

Here, pp is the positive share of the sampled page stock; λ\lambda is an annual replacement, rewriting, and dilution hazard; and q(t)q(t) is the positive share of incoming or substantially refreshed content. It rises linearly from q0q_0 to qTq_T. Starting from p0=0.096p_0=0.096, the endpoint is:

p(T)=q0+(p0−q0)e−λT+qT−q0T[T−1−e−λTλ].p(T)=q_0+(p_0-q_0)e^{-\lambda T}+\frac{q_T-q_0}{T}\left[T-\frac{1-e^{-\lambda T}}{\lambda}\right].

The horizon is a judgmental 4.25 years, favoring a later qualifying snapshot while allowing snapshot timing uncertainty in the predictive spreads. The assumptions below are my forecast specification, informed by the evidence above—not empirically identified parameter estimates (baseline input).

ScenarioWeightAnnual hazardInitial flow positive shareTerminal flow positive shareEndpoint meanPredictive SD, percentage points
Slow diffusion and persistent legacy24%0.06535%50%17.62%6.5
Continued diffusion56%0.15040%70%31.49%10.5
Faster rewriting and proliferation18%0.28050%90%54.32%16.0
Unmodeled structural or measurement break2%———Uniform over 0%–100%—

The central scenario initially grows by 4.56 percentage points annually, close to the recent fitted slope of 4.88 points. The slow case requires a slowdown; the fast case requires acceleration. Persistent pages and the article plateau justify substantial slow-case weight. Integrated editing, cheap page creation, and improved mixed-authorship detection justify retaining a meaningful fast-growth tail (historical input; publishing mechanism; measurement mechanism).

For each structural scenario, I use a bounded Beta distribution matching its endpoint mean and predictive standard deviation. Those spreads cover adoption, renewal, crawl mix, detector redesign, and snapshot timing. They deliberately preserve the disagreement between reasonable modeling approaches rather than narrowing around one fitted curve. The remaining tail reserve protects against structural changes those models omit.

The resulting distribution has a median of 30%, a mean of 32.8%, a central 90% interval of 11%–67%, and a 3% probability of a result at or below 10% (forecast specification above). I assign 70% probability to fine-precision reporting and 30% to whole-percentage-only reporting. These are reporting assumptions, not permission to ignore an available unrounded result. Rounding is applied before bucket allocation, giving integer outcomes the appropriate extra mass. The supplied code returns the complete normalized distribution without rounding its probabilities.

What's non-obvious

A plateau in new-content adoption does not imply a plateau in the web stock. In the central model, holding incoming-content positivity fixed at 40%, rather than increasing it, still raises the all-page share to about 24% as old material is displaced or rewritten. That is my calculation from the baseline and stock model. Persistent URLs are not a protected reservoir of unchanged human prose: AI can substantially rewrite a page without changing its address.

Higher headline estimates often concern a selected denominator. A September 30, 2026 preprint reports 31.1% AI-labeled tokens in August’s FineWeb-filtered text—not raw English-page prevalence. Its July filtering audit used 10,000 documents from 32 WARC files and retained AI-labeled documents 2.3 times as often as human-labeled ones; the audit also precedes English filtering, so its raw counts are not a direct English-page replication. This prevents rebasing the forecast to that headline while still showing how extraction and filtering can move the measured share (study and Appendix A). (arxiv.org)

Uncertainties

The largest missing measurement is the future weight of unchanged older text. URL survival, crawler discovery, and publication dates do not identify that quantity. A representative longitudinal panel measuring substantive text changes would sharpen the renewal hazard much more than another article-only estimate (accessibility study; publication-date limitations).

Detector agreement is not ground truth. Pew’s paired check covered 62,370 pages from seven crawls, with 96% agreement and a kappa of 0.61. Its early positive readings likely include false positives. No verified bridge establishes how a successor detector would classify the same representative web sample against independently known substantial editing. That gap prevents a fixed upward adjustment for improved detection (Pew validation; successor-model evaluation).

The evidence cutoff is the supplied October 5, 2026 date. The current positive share is not a hard floor: this is a snapshot proportion, and valid reclassification, page loss, or crawl-composition changes can lower it. The distribution is conditional on a qualifying numerical result under the supplied resolution rules; failure to obtain a qualifying study by December 31, 2032 means annulment, not zero. No prediction-market prices or public forecast aggregates enter the estimate (forecast interpretation and specification).

Sources

  1. Domain Expert Search · mcp

    Found 7 domain experts for 'AI text detection, mixed human-AI authorship, Common Crawl sampling bias and measurement comparability':

  2. Claude Code · e2b

    Job coding_whiz_job_53a26470bb done after 328936ms.

  3. How Much of the Internet Is Written With AI? | Pew Research Center · openai
  4. Methodology | Pew Research Center · openai
  5. pewresearch.org · tool
  6. pewresearch.org · tool
  7. Domain Expert Research Task · mcp

    Job domain_expert_research_task_fd65196f14 done after 329474ms.

  8. Pangram 4 Technical Report · openai
  9. arxiv.org · tool
  10. Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift · openai
  11. arxiv.org · tool
  12. pangram.com · tool
  13. arxiv.org · tool
  14. arxiv.org · tool
  15. pangram.com · tool
  16. The Impact of AI-Generated Text on the Internet · anthropic
  17. arxiv.org · tool
  18. AI Now Writes as Many Online Articles as Humans — Five Percent · openai
  19. Link Rot and Digital Decay on Government, News and Other Webpages | Pew Research Center · openai
  20. What's new on WordPress.com: AI for all paid plans, more · openai
  21. How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text · openai
  22. Epoch · mcp

    Poll 'mar_2026' (2026-03): 59 rows.

  23. epoch.ai · tool
  24. News · mcp

    Found 20 merged articles (asknews: 10, perigon: 10, both: 0).

  25. wpvip.com · tool
  26. techrepublic.com · tool
  27. trendhunter.com · tool
  28. techrepublic.com · tool
  29. digiday.com · tool
  30. cnbc.com · tool
  31. contently.com · tool
  32. reuters.com · tool
  33. seroundtable.com · tool
  34. bizjournals.com · tool
  35. tag24.com · tool
  36. forbes.com · tool
  37. medium.com · tool
  38. forbes.com · tool
  39. zdnet.com · tool
  40. forbes.com · tool
  41. trendhunter.com · tool
  42. thestar.com.my · tool
  43. businessinsider.com · tool
  44. channelnewsasia.com · tool
  45. Anthropic Economic Index Data · mcp

    ANTHROPIC ECONOMIC INDEX - 6 release(s), oldest first

  46. Cloudflare Radar · mcp

    HTTP traffic share by bot_class (%)

  47. ai-on-the-internet.github.io · tool
  48. The Widespread Adoption of Large Language Model-Assisted Writing Across Society · openai
  49. pmc.ncbi.nlm.nih.gov · tool
  50. arxiv.org · tool

Question Details

Description

This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.

Resolution Criteria

Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.

Fine Print

“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.