Forecast report
What percent of webpages will be significantly written or edited by AI in 2030?
Forecast
Median forecast: 29; 80% interval: 14 to 57.5.
Distribution
Analysis
TL;DR
My median forecast is 29% of all qualifying English-language webpages in the resolving snapshot (forecast model below). The central 80% forecast interval is 14%–57.5%, with a 15% probability of a result above 50% (distribution calculation). This forecasts the study-classified share of the whole sampled web, conditional on a numerical resolution—not the share of newly published articles or AI-generated words.
Context
The correct starting point is the July 18, 2026 all-page reading of 9.60%, not the 35.09% reading for dated pages published after ChatGPT’s release (Pew baseline, published August 20, 2026). Only about 10%–15% of sampled pages had identifiable publication dates, and that subset was nonrandom (Pew methodology). Using the larger number as the baseline would substantially overstate the quantity being forecast.
The target is a page-level text classification, not proven authorship. The sampled web also excludes some material available to human readers: paywalled and login-required pages are underrepresented (Pew methodology). The numerical distribution is conditional on a qualifying study; annulment is not assigned to any percentage bucket.
Evidence
The historical backbone consists of 49 crawl samples, each containing 10,000 English-language pages, or 490,000 page-snapshot observations (Pew sampling methodology). Below is the complete all-page history from the published interactive chart. Dates identify snapshots, not original publication dates; values are percentages of pages classified as showing meaningful AI writing or editing (Pew chart data, August 20, 2026 vintage). Read each row from left to right.
| Snapshot | % | Snapshot | % | Snapshot | % |
|---|---|---|---|---|---|
| 2021-01-27 | 1.18 | 2021-03-03 | 1.13 | 2021-04-20 | 1.10 |
| 2021-05-17 | 1.12 | 2021-06-23 | 1.08 | 2021-08-01 | 1.11 |
| 2021-09-19 | 1.07 | 2021-10-20 | 1.13 | 2021-11-30 | 1.09 |
| 2022-01-24 | 1.10 | 2022-05-25 | 1.04 | 2022-07-06 | 1.04 |
| 2022-08-16 | 1.02 | 2022-09-29 | 1.07 | 2022-12-08 | 1.15 |
| 2023-02-04 | 1.35 | 2023-03-20 | 1.53 | 2023-06-09 | 1.89 |
| 2023-09-21 | 2.30 | 2023-12-10 | 2.78 | 2024-03-01 | 3.22 |
| 2024-04-13 | 3.55 | 2024-05-18 | 3.49 | 2024-06-24 | 3.49 |
| 2024-07-24 | 3.40 | 2024-08-04 | 3.80 | 2024-09-20 | 3.83 |
| 2024-10-12 | 4.34 | 2024-11-02 | 4.51 | 2024-12-10 | 4.92 |
| 2025-01-22 | 4.86 | 2025-02-06 | 4.92 | 2025-03-23 | 4.98 |
| 2025-04-22 | 5.07 | 2025-05-12 | 4.93 | 2025-06-22 | 4.86 |
| 2025-07-14 | 4.78 | 2025-08-13 | 5.24 | 2025-09-12 | 5.49 |
| 2025-10-15 | 6.01 | 2025-11-15 | 6.29 | 2025-12-11 | 6.64 |
| 2026-01-25 | 6.92 | 2026-02-19 | 7.28 | 2026-03-15 | 8.09 |
| 2026-04-20 | 8.58 | 2026-05-19 | 8.86 | 2026-06-06 | 8.98 |
| 2026-07-18 | 9.60 | — | — | — | — |
I fitted calendar-time linear trends and anchored their projections to the last observation, rather than to a fitted endpoint that misses recent acceleration. These calculations give materially different answers depending on the historical window (input observations):
| Estimation window, ending July 18, 2026 | Observations | Slope, percentage points/year | Projection to September 30, 2030 |
|---|---|---|---|
| From December 8, 2022 | 35 | 2.110 | 18.5% |
| From July 24, 2024 | 25 | 2.715 | 21.0% |
| From August 13, 2025 | 12 | 4.819 | 29.8% |
These are extrapolation checks, not statistical prediction intervals. The long-window estimate preserves early slow adoption; the short-window estimate preserves recent acceleration. Neither establishes a future saturation level. I preserve that disagreement rather than treating several fits to the same data as independent confirmation.
Legacy content supplies real inertia. A study published May 17, 2024 checked nearly one million URLs drawn from 2013–2023 snapshots in October 2023. Only 38% of the oldest snapshot’s URLs were confirmed inaccessible after a decade (Pew digital-decay study). But URL survival does not prove text survival: an old page can be substantially rewritten without changing its address.
Nor does crawler discovery establish publication. The September 2026 crawl, collected September 4–17 and announced September 19, contained about 2.17 billion captures, including 587.2 million previously unvisited URLs—roughly 27.1% by division of the rounded counts (Common Crawl release). Those are all-language capture counts. A newly discovered URL can contain old text, while a repeatedly captured URL can contain new text. I do not insert this discovery percentage into an annual replacement model.
The incoming-content evidence is much more AI-heavy. A May 2026 analysis sampled 55,400 English-language articles and listicles with publication dates spanning January 2020–March 2026, requiring article markup and at least 100 words. It averaged three detectors (Graphite methodology). Its complete publication-cohort history follows; these are retrospective classifications of article cohorts, not comparable all-page snapshots (published source data).
| Publication year | Q1 primarily AI, % | Q2, % | Q3, % | Q4, % |
|---|---|---|---|---|
| 2020 | 0.97 | 0.90 | 0.95 | 0.64 |
| 2021 | 1.04 | 1.23 | 0.89 | 1.18 |
| 2022 | 1.54 | 2.37 | 3.85 | 4.60 |
| 2023 | 14.02 | 25.14 | 33.69 | 35.92 |
| 2024 | 38.23 | 37.22 | 41.81 | 47.04 |
| 2025 | 49.61 | 44.58 | 49.22 | 50.90 |
| 2026 | 49.94 | — | — | — |
The recent plateau argues against automatic exponential growth. But this sample excludes many page types, and its tests did not validate naturally occurring workflows in which people rewrite AI drafts (Graphite limitations). I use it to bound plausible incoming adoption, not to replace the broad-web baseline.
Publishing infrastructure provides a stronger leading signal than generic chatbot-use surveys. WordPress.com announced direct agent creation, editing and publication capabilities on March 20, 2026, with the announcement updated May 15 (agent publishing announcement). Its May 8 changelog expanded the opt-in editor assistant to all current paid plans (editor rollout). This is WordPress.com, not the entire self-hosted ecosystem. My inference is that native assistance makes substantive editing of existing pages cheaper, creating a growth path beyond new articles. Availability alone does not establish actual publication rates.
Search incentives provide a brake, not a ban. Google targets mass-produced content intended to manipulate rankings without adding value, including AI-generated material; it does not prohibit useful AI assistance (Google spam policy). Lower search visibility can reduce publishing returns without removing those pages from a broad archive.
Measurement can move the answer in either direction. In a July 29, 2026 vendor report, the same 4,826 synthetically and substantially edited texts received AI-or-Mixed classifications in 34.7% of cases under Pangram 3.3.2 versus 72.2% under Pangram 4 (technical report, Table 6). This demonstrates method sensitivity, not a correction factor for random webpages. Conversely, external research dated September 25 demonstrated detector evasion through agent-assembled base-model text, with experiments using 100 examples per benchmark and substantially higher generation costs (HALO study). Neither experiment estimates ordinary-web prevalence. Together they justify wide, two-sided measurement uncertainty.
My quantitative model treats the target as a stock renewed by new or substantially rewritten content:
Here is the classified share of the sampled stock, is the comparable classified share of incoming or substantially refreshed text, and is an effective renewal hazard per year. Renewal includes new publication, major edits, deletion and changing crawl inclusion. Major legacy editing is included once, not added again as a separate conversion channel. The code starts at the measured baseline and lets change linearly over the horizon.
The following are my forecasting assumptions, not measured turnover rates or fitted adoption ceilings. Timing is centered on a late-year snapshot; the outcome spreads cover earlier or later qualifying snapshots (reproducible calculation).
| Scenario | Weight | Renewal hazard/year | Incoming classified share, start → finish | Calculated outcome center | Within-scenario SD, percentage points |
|---|---|---|---|---|---|
| Slow conversion and restrained adoption | 23% | 0.075 | 35% → 43% | 17.6% | 7 |
| Continued diffusion | 52% | 0.140 | 40% → 65% | 28.9% | 10 |
| Rapid publishing and inventory rewriting | 20% | 0.270 | 45% → 85% | 49.5% | 14 |
| Transformative automation | 3% | 0.600 | 50% → 95% | 75.2% | 13 |
The central case produces an initial increase of about 4.3 percentage points annually, close to the recent trend calculation. Its rising incoming share reflects wider substantive assistance; its moderate renewal rate preserves substantial legacy content. The slow case retains the plateau and persistence evidence. The fast cases retain the possibility of a structural change in publishing economics rather than assuming today’s constraints last indefinitely.
Each scenario becomes a bounded beta distribution matching its center and stated standard deviation. A further 2% weight is uniform over the allowed range for unmodeled but qualifying regimes. These widths cover parameter, detector, crawl-composition and timing uncertainty. They deliberately do not shrink merely because the historical sample is large.
The code also models reporting precision: 60% weight on hundredths of a percentage point, 15% on tenths, and 25% on whole percentages. These are judgmental reporting assumptions. Whole-number reporting puts more mass in integer-ending buckets. The resulting distribution has a median of 29%, a mean of 32.4%, and a central 90% interval of 11%–68%; the right tail raises the mean above the median (distribution calculation).
What's non-obvious
A quality filter can increase the apparent AI share. A September 30, 2026 preprint audited 10,000 raw crawl documents: FineWeb retained 29.3% of AI-labeled documents versus 12.8% of human-labeled documents. Even among English documents, its quality heuristics retained about 64% versus 39% (filtering audit). The paper is vendor-linked, and its labels are not ground truth. Still, it identifies a concrete selection effect: cleaner-looking training corpora need not represent the whole sampled web. Filtered token estimates therefore cannot simply replace a raw, page-weighted baseline.
Crawler access creates another selection effect. Cloudflare’s managed robots file explicitly blocks CCBot (managed-robots documentation), and Common Crawl respects those exclusions (crawler FAQ). If human-heavy publishers opt out more often than automated publishers, the archive’s classified share can rise without an equal rise across the public web; the reverse is also possible. I could not quantify that difference. This adds uncertainty about the denominator, not evidence for a fixed upward adjustment.
Uncertainties
The largest missing input is a representative panel measuring how often existing English-language pages receive substantial AI edits. URL disappearance and first crawler discovery do not answer that question. A second gap is independent calibration on naturally mixed webpages: the synthetic mixed-edit gains do not establish how many randomly sampled pages a newer detector would reclassify (Pangram evaluation). A third is changing crawl access and composition, including whether a future study can retain a sufficiently representative denominator.
Sampling noise is small beside these gaps. At the baseline prevalence and 10,000 independent pages, the binomial standard error is about 0.3 percentage points—a calculation using the reported sample size (Pew methodology). That does not cover correlated pages, classifier error or a changing sampling frame. I therefore keep a broad forecast, including a 4% probability of an outcome at or below 10%, rather than treating the current reading as a hard lower bound (forecast distribution). The precise probability array represents numerical precision in this model, not equally precise knowledge of the future.
Sources
- Domain Expert Search · mcp
Found 7 domain experts for 'AI text authorship measurement Pangram mixed human-AI editing detection research systematic webpage prevalence':
- Epoch · mcp
Poll 'mar_2026' (2026-03): 59 rows (showing 10).
- epoch.ai · tool
- News · mcp
Found 6 merged articles (asknews: 3, perigon: 3, both: 0).
- businessinsider.jp · tool
- dailykos.com · tool
- seosandwitch.com · tool
- postandcourier.com · tool
- statista.com · tool
- finance.yahoo.com · tool
- Claude Code · e2b
Job coding_whiz_job_b2be430d2d done after 271421ms.
- How Much of the Internet Is Written With AI? | Pew Research Center · openai
- Methodology | Pew Research Center · openai
- Common Crawl - Blog - September 2026 Crawl Archive Now Available · openai
- Domain Expert Research Task · mcp
Job domain_expert_research_task_a05d26855e done after 208424ms.
- Pangram 4 Technical Report · openai
- Agents Can Use Base Models to Evade AI Detection · openai
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text · openai
- The Impact of AI-Generated Text on the Internet · openai
- arxiv.org · tool
- Anthropic Economic Index Data · mcp
No rows in release 2026-06-26 for breakdown='task', granularity='broadest', geo_id='GLOBAL', query='writing'. Widen the query or try geo_id='GLOBAL'.
- Openalex · mcp
Found 988 works matching 'AI generated text web Common Crawl prevalence detection'. Showing 20:
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- doi.org · tool
- blog.hubspot.com · tool
- computerworld.com · tool
- medium.com · tool
- searchenginejournal.com · tool
- visualcapitalist.com · tool
- bizjournals.com · tool
- seosandwitch.com · tool
- techrepublic.com · tool
Question Details
Description
This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.
Resolution Criteria
Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.
Fine Print
“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.