Forecast report
What percent of webpages will be significantly written or edited by AI in 2030?
Forecast
Median forecast: 29.5; 80% interval: 14 to 56.
Distribution
Analysis
TL;DR
My central forecast is 30% of all sampled webpages, with an expected value of 32.7% (forecast calculation). The central 90% predictive interval is 11%–67%, and the probability of exceeding 50% is 15% (forecast calculation). This measures the whole webpage stock, not newly published articles, and is conditional on a qualifying numerical resolution (sampling method; resolution interpretation).
Context
The direct baseline is 9.60% in the July 2026 snapshot, rounded to 10% in Pew’s August 20, 2026 report. The more precise value is present in the article’s embedded chart data (Pew report). This forecast uses information available on October 3, 2026 (forecast vintage).
The target counts substantial AI editing as well as authorship. It does not require AI to have generated most of a page, and the detector’s 0.2 cutoff is not a verified fraction of AI-written words (Pew methodology). I allow future detector changes that preserve this concept, rather than forecasting the output of a frozen classifier (forecast definition).
Evidence
The historical backbone is Pew’s random sample of 10,000 English-language pages from each of 49 crawls, totaling 490,000 pages (methodology). Below is the complete all-pages chart history, spanning January 27, 2021–July 18, 2026, at its published August 20, 2026 vintage. Each value is a percentage of pages classified positive, not a percentage of words (primary chart data; data table).
| Crawl date | AI-positive % | Crawl date | AI-positive % |
|---|---|---|---|
| 2021-01-27 | 1.18 | 2021-03-03 | 1.13 |
| 2021-04-20 | 1.10 | 2021-05-17 | 1.12 |
| 2021-06-23 | 1.08 | 2021-08-01 | 1.11 |
| 2021-09-19 | 1.07 | 2021-10-20 | 1.13 |
| 2021-11-30 | 1.09 | 2022-01-24 | 1.10 |
| 2022-05-25 | 1.04 | 2022-07-06 | 1.04 |
| 2022-08-16 | 1.02 | 2022-09-29 | 1.07 |
| 2022-12-08 | 1.15 | 2023-02-04 | 1.35 |
| 2023-03-20 | 1.53 | 2023-06-09 | 1.89 |
| 2023-09-21 | 2.30 | 2023-12-10 | 2.78 |
| 2024-03-01 | 3.22 | 2024-04-13 | 3.55 |
| 2024-05-18 | 3.49 | 2024-06-24 | 3.49 |
| 2024-07-24 | 3.40 | 2024-08-04 | 3.80 |
| 2024-09-20 | 3.83 | 2024-10-12 | 4.34 |
| 2024-11-02 | 4.51 | 2024-12-10 | 4.92 |
| 2025-01-22 | 4.86 | 2025-02-06 | 4.92 |
| 2025-03-23 | 4.98 | 2025-04-22 | 5.07 |
| 2025-05-12 | 4.93 | 2025-06-22 | 4.86 |
| 2025-07-14 | 4.78 | 2025-08-13 | 5.24 |
| 2025-09-12 | 5.49 | 2025-10-15 | 6.01 |
| 2025-11-15 | 6.29 | 2025-12-11 | 6.64 |
| 2026-01-25 | 6.92 | 2026-02-19 | 7.28 |
| 2026-03-15 | 8.09 | 2026-04-20 | 8.58 |
| 2026-05-19 | 8.86 | 2026-06-06 | 8.98 |
| 2026-07-18 | 9.60 | — | — |
The early positives are a detector-background signal, not evidence that substantial AI authorship was already common before ChatGPT. Pew explicitly warns about false positives in those samples (methodology).
The full series shows why a single trend line gives false precision. A linear fit to the 35 observations from December 2022–July 2026 projects 17.5% for December 31, 2030; fitting the 13 observations from July 2025–July 2026 gives 30.9%. On the full post-release window, bounded logistic and Gompertz curves with the same assumed 40% ceiling and 1.05% background give 34.5% and 25.4%, respectively. These are my sensitivity calculations, not published forecasts; the history does not identify a saturation ceiling (calculations and specifications).
Legacy content is durable. Pew’s May 17, 2024 study sampled nearly one million pages in annual cohorts covering 2013–2023, approximately 90,000 per cohort, and checked accessibility in October 2023. Of the oldest cohort, 38% was inaccessible after a decade. Accessibility does not tell us whether surviving text was rewritten (digital-decay study). A July 15, 2026 study by Common Crawl engineers, using its 2020–2025 archives, also found that a homogeneous turnover model failed: a persistent domain core and a changing outer population fit better. That supports heterogeneous renewal, not one physical half-life for all webpages (crawl-persistence research).
New-content measurements are much higher, but their populations differ. Pew reports 35% for dated, post-ChatGPT pages in July 2026; publication dates exist on only 10%–15% of sampled pages, and this subset is nonrandom (methodology). Graphite’s May 15, 2026 study reports 49.9% primarily AI-generated articles in the first quarter of 2026. Its overall sample contains 55,400 English-language articles/listicles dated January 2020–March 2026, selected for article markup and at least 100 words; quarterly sample counts are not given. This commercial study informs incoming article prevalence, not all-page prevalence (study and selection rules).
A September 30, 2026 University of Maryland/Pangram preprint reinforces the denominator problem. Its 310,000-document sample spans January 2021–August 2026 and reports a 31.1% AI-labelled token share for August after FineWeb filtering. A separate 10,000-document audit found AI-labelled documents surviving those filters 2.3 times as often as human-labelled documents. Filtered tokens therefore cannot be converted into raw page prevalence by assuming equal document lengths or equal retention (preprint and methods). I use this as evidence against mechanically combining headline percentages.
Publishing tools create an editing channel beyond new articles. WordPress’s May 20, 2026 release put an AI client into its core, enabling model connections and publishing workflows. Availability is not uptake, but it reduces friction (release documentation). Google’s guidance, updated October 1, 2026, targets mass generation without user value rather than AI assistance itself. I read these incentives as favoring reviewed, mixed-authorship content over unlimited low-value generation (search guidance).
Measurement can move in either direction. Pangram’s vendor-authored July 29, 2026 evaluation of 14,990 student essays subjected to substantial-editing instructions classified 21% as Mixed or AI with its older commercial model, versus 58.6% with its newer model. That is not a random-web validation or a test of Pew’s exact open detector, so I apply no fixed correction (technical report). Conversely, independent research published May 19, 2026 found base-model continuations much more likely to be judged human than instruction-tuned continuations. A changed writing distribution can weaken detection even without deliberate evasion (independent study).
My main calculation treats prevalence as renewal of a mixed-age stock:
Here, is the measured positive fraction, is the positive fraction among incoming or replacement pages, is the effective renewal hazard, is additional substantial editing of surviving nonpositive pages, and is additional loss of positive status through differential disappearance or detectability. The hazards have units of inverse years. Incoming prevalence rises linearly within each scenario. The starting fraction is 0.096 and the reference horizon is 4.2 years; uncertainty about the resolving snapshot’s date is included in the predictive widths (model calculation).
The following rates and weights are my forward-looking assumptions, not measured turnover rates or a fitted posterior. Renewal includes changing crawl composition as well as physical replacement (scenario calculation).
| Scenario | Weight | Renewal/year | Extra editing/year | Extra loss/year | Incoming share: start → end | Model center |
|---|---|---|---|---|---|---|
| Slow accumulation | 25% | 0.07 | 0.003 | 0.014 | 30% → 40% | 16% |
| Continued adoption | 50.0% | 0.13 | 0.012 | 0.006 | 38% → 62% | 30% |
| Fast expansion and rewriting | 20% | 0.23 | 0.025 | 0.003 | 42% → 85% | 49% |
| Transformative expansion | 5% | 0.50 | 0.060 | 0.008 | 55% → 95% | 76% |
The central scenario reconciles recent acceleration with surviving legacy content. The slow scenario retains the long-window trend and stricter measurement. The fast scenarios retain the possibility of much larger publishing volumes, widespread rewriting, and better recognition of mixed authorship. Their weights preserve uncertainty across competing models rather than treating shared evidence as independent confirmation (scenario calculation).
Each scenario has a beta distribution on the bounded fraction scale, with parameters and , where is its model center and its concentration. Concentrations are 35, 35, 24 and 18, giving within-scenario standard deviations of roughly 6, 8, 10 and 10 percentage points. These are subjective predictive widths, not sampling errors. I also assign 60.0% probability to reporting in hundredths, 20% to tenths, and 20% to whole percentages. The probability array preserves the extra mass in buckets containing rounded values (distribution calculation).
What's non-obvious
Old content buffers the all-page share, but it does not form a fixed denominator. Common Crawl says it revisits slowly changing pages less often and concentrates on newly discovered or changing pages. A snapshot can therefore become more AI-heavy without most surviving old pages being rewritten (crawler’s coverage explanation, July 20, 2026). “New to the archive” also does not mean “newly published”; the September 19, 2026 release explicitly describes previously unvisited URLs, not publication dates (archive release).
The apparent precision of a large sample is misleading. At the baseline, an independent 10,000-page sample has a binomial standard error of only 0.3 percentage points, but this excludes clustering, coverage and classification error (calculation). Better recognition of substantial editing can raise the measured result, while more human-like generation can lower it. Neither effect is resolved by agreement between two detectors (editing benchmark; detector-shift evidence).
Uncertainties
The largest missing measurement is effective renewal of English webpage text. Repeated observations of the same URLs, with text changes and genuine publication dates, would distinguish old pages being edited from new pages entering and crawler priorities changing. Existing persistence research does not identify that rate (coverage research).
The second gap is ground-truth authorship. An audited random web sample with known drafting and editing histories would sharpen detector calibration. The available mixed-authorship benchmarks establish mechanisms, not a web-wide percentage-point correction, and a selected dated-page sample does not identify the web’s age distribution (editing benchmark; sampling limitations).
Finally, the qualifying study, snapshot date and future classifier remain unknown. This uncertainty belongs in the wide distribution, not in a claim that prevalence must exceed the baseline. Under the client’s rules, absence of a comparable study by December 31, 2032 means annulment; it is not a zero-percent outcome. The numerical distribution is conditional on resolution (resolution interpretation and calculation).
Sources
- Domain Expert Search · mcp
Found 7 domain experts for 'AI-generated web text measurement, Common Crawl sample representativeness, AI detection and mixed-authorship validation':
- Domain Expert Research Task · mcp
Job domain_expert_research_task_262d0a2f29 done after 230022ms.
- Methodology | Pew Research Center · openai
- arxiv.org · tool
- How Much of the Internet Is Written With AI? | Pew Research Center · openai
- huggingface.co · tool
- Pangram 4 Technical Report · openai
- arxiv.org · tool
- arxiv.org · tool
- [2605.19516] Base Models Look Human To AI Detectors · openai
- arxiv.org · tool
- arxiv.org · tool
- proceedings.mlr.press · tool
- Claude Code · e2b
Job coding_whiz_job_1d3c1473da done after 420699ms.
- Common Crawl - Blog - July 2026 Crawl Archive Now Available · openai
- Common Crawl - Blog - September 2026 Crawl Archive Now Available · anthropic
- Common Crawl - Blog - Measuring Crawled Coverage of a Website in Common Crawl · openai
- arXiv · mcp
arXiv ID: 2604.26965v1
- arxiv.org · tool
- arxiv.org · tool
- Epoch · mcp
Poll 'mar_2026' (2026-03): 59 rows.
- epoch.ai · tool
- huggingface.co · tool
- arxiv.org · tool
- github.com · tool
- Statistics of Common Crawl Monthly Archives by commoncrawl · openai
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- arxiv.org · tool
- link.springer.com · tool
- arxiv.org · tool
- Link Rot and Digital Decay on Government, News and Other Webpages | Pew Research Center · openai
- mlanthology.org · tool
- researchonline.jcu.edu.au · tool
- The Impact of AI-Generated Text on the Internet · anthropic
- Introducing Open Pangram | Pangram · openai
- AI Now Writes as Many Online Articles as Humans — Five Percent · openai
- ahrefs.com · tool
- ahrefs.com · tool
- typeform.com · tool
- clutch.co · tool
Question Details
Description
This question asks for the percentage of English-language webpages in a representative 2030 snapshot of the publicly accessible web that show significant signs of having been written or substantially edited by AI, using a methodology comparable to Pew Research Center's 2026 study “How Much of the Internet Is Written With AI?” As background, Pew sampled 10,000 English-language webpages from each Common Crawl snapshot in its 2021–2026 analysis and applied Pangram's open-weight AI-detection model to the pages' body text. Pew classified pages scoring at least 0.2 as showing meaningful signs of AI authorship or editing. In its July 2026 sample, 10% of all webpages showed significant signs of AI authorship; among pages with identifiable publication dates after ChatGPT's November 30, 2022 release, the share was about 35%. The question concerns the former, all-webpages measure rather than the subset restricted to recently published pages. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) The target is the percentage reported for a web snapshot collected during calendar year 2030, preferably using Common Crawl or a successor web archive and a large random sample of English-language webpages, with an AI-authorship detector and classification threshold designed to measure the same underlying concept as Pew's 2026 analysis. Because AI-detection technology will likely change by 2030, methodological comparability of the measured concept is more important than requiring use of exactly the same model version.
Resolution Criteria
Resolve to the percentage of sampled English-language webpages that a qualifying 2030 study classifies as showing significant or meaningful signs of AI authorship or substantial AI editing. The primary resolution source will be a Pew Research Center study published in or after 2030 that updates its August 20, 2026 analysis using a web snapshot from calendar year 2030 and reports the corresponding all-webpages percentage. If Pew publishes multiple qualifying estimates based on 2030 snapshots, use the estimate corresponding to the latest 2030 snapshot. Use the study's reported unrounded value if available; otherwise use its reported rounded percentage. A qualifying update should be methodologically comparable in its target quantity to Pew's 2026 study: it should sample English-language webpages from a broad archive or crawl of the publicly accessible web and use systematic text analysis to estimate the share showing substantial AI authorship or editing. It need not use the exact 2026 Pangram model or the exact 0.2 threshold if researchers change methods to maintain or improve validity as AI and detection methods evolve. Pew's 2026 methodology used random samples from Common Crawl and treated an Open Pangram score of at least 0.2 as meaningful evidence of AI authorship/editing. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The denominator is all qualifying English-language webpages in the sampled 2030 web snapshot, not only webpages first published in 2030, pages published after ChatGPT's release, newly published pages, or pages with detectable publication dates. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/)) If Pew does not publish a qualifying result by December 31, 2032, use a peer-reviewed study or a study from another established research institution that applies a substantially comparable methodology to a broad, representative sample of English-language webpages from 2030. If no sufficiently comparable study using a 2030 web snapshot is available by that date, annul the question rather than extrapolating from another year.
Fine Print
“Significantly written or edited by AI” refers to the study's page-level classification of meaningful/substantial AI authorship or editing, not proof that AI generated most or all of a page. AI detectors are probabilistic and can produce false positives and false negatives; Pew explicitly cautions that individual classifications are not definitive, while using aggregate results to track prevalence across large samples. ([pewresearch.org](https://www.pewresearch.org/data-labs/2026/08/20/methodology-ai-content/)) The target concerns text content. AI-generated images, audio, video, code, page layouts, recommendation systems, or other non-textual uses of AI do not by themselves make a webpage count as AI-written or AI-edited. If an otherwise qualifying study reports several estimates under alternative detectors or specifications without identifying a preferred headline estimate, use the estimate its authors designate as their primary result. If no primary estimate is designated and the alternatives cannot be reconciled into one clearly comparable headline measure, the question should be annulled rather than resolving via an arbitrary choice.