Forecast report
Will a Chinese model hold the AA Intelligence Index and the Vals Index simultaneously at any point between now and the end of 2028?
Forecast
P(Yes): 28.5%; P(No): 71.5%.
Distribution
Analysis
TL;DR
I assign a 28% probability of YES. The strongest currently evaluated Chinese model trails the leaders on both required indices (Artificial Analysis, Vals). The opportunity is one brief overlapping lead, not permanent Chinese dominance.
Context
I use the client's September 30, 2026, 08:34 UTC cutoff and the full eligibility window, beginning July 22, 2026, at 22:46 UTC and ending December 31, 2028. I require the same identifiable Chinese model to hold first place on both overall boards during an overlapping interval. Official joint leadership counts; overlapping error bars alone do not. Ordinary differences in reasoning-effort settings do not change the underlying model's identity.
The current leaders are Claude Opus 5.5 on AA Intelligence Index v4.3.2 and Claude Sonnet 5.5 on Vals Index v2.1. Xiaomi's MiMo-V2.6-Pro is the strongest evaluated Chinese contender on both. I found no verified qualifying historical episode, but the archives do not cover every moment (AA rankings, Vals rankings, updated September 29, historical AA snapshot).
Evidence
The historical backbone favors continued Chinese catch-up without guaranteed overtaking. Epoch AI's January 2, 2026 analysis reported that every frontier model in its ECI record since 2023 had been developed in the United States. Across its 2023–late-2025 coverage, Chinese models lagged by seven months on average, with a four-to-fourteen-month range. This is an older, different index. The report does not supply an independent-trial sample size, so I use it as evidence of persistent lag, not as a fitted probability of this event (Epoch analysis).
Broader rankings give a different picture. Stanford's report describes repeated US–Chinese changes at the top of performance rankings since early 2025. Its discussion relies on broader measures, including human preference, rather than these two professional-agent indices. That rejects an extremely pessimistic prior, but does not establish leadership on the required boards (Stanford technical-performance review).
The complete set of recovered archived overall-leader observations used here is below. These are eight AA captures and six Vals index-bearing page states, spanning July 22–September 30. They are sparse, dependent observations—not fourteen independent trials. The two opening-day captures precede question creation. Scores are contemporary archived values, not current scores assigned retrospectively to older releases.
| Index | Capture timestamp, UTC | Reported overall leader | Score |
|---|---|---|---|
| Vals | 2026-07-22 16:48:28 — pre-window | Claude Fable 5 | 75.14% |
| AA | 2026-07-22 17:37:08 — pre-window | Claude Fable 5 | 60 points |
| Vals | 2026-07-26 20:22:54 | Claude Fable 5 | 75.14% |
| AA | 2026-07-31 04:59:10 | Claude Opus 5 | 61 points |
| AA | 2026-08-06 07:19:46 | Claude Opus 5 | 61 points |
| Vals | 2026-08-15 10:01:32 | Claude Opus 5 | 67.21% |
| AA | 2026-08-19 02:53:23 | Claude Opus 5 | 63 points |
| AA | 2026-08-20 12:20:47 | Claude Opus 5 | 63 points |
| AA | 2026-08-21 02:48:50 | Claude Opus 5 | 63 points |
| AA | 2026-08-31 18:51:03 | Claude Opus 5 | 63 points |
| Vals | 2026-09-09 12:46:17 | Claude Fable 5.1 | 68.83% |
| Vals | 2026-09-21 06:55:33 | Claude Fable 5.1 | 68.83% |
| Vals | 2026-09-29 20:25:37 | Claude Sonnet 5.5 | 67.04% |
| AA | 2026-09-30 03:05:14 | Claude Opus 5.5 | 58 points |
The strongest positive historical evidence is a near-miss. AA's July 17 report put Kimi K3 at 57, behind Fable 5 at 60. The Vals page state dated July 23, captured July 26, put Kimi at 74.70%, only 0.44 percentage points behind Fable's 75.14%. Kimi also performed strongly on private knowledge-work evaluations. Chinese proximity therefore extended beyond public mathematics and coding tests (AA's Kimi evaluation, archived Vals results).
Today's deficits are substantially larger. The September 30 AA data give Opus 5.5 maximum effort 57.6 points and MiMo Pro 46.3, an 11.3-point difference; public pages round those scores. Vals's September 29-updated board gives Sonnet 5.5 67.04% and MiMo 55.20%, an 11.84-percentage-point difference. MiMo ranks eleventh among thirty Vals entries. AA points and Vals weighted accuracy are different units and should not be averaged (AA leader, AA MiMo results, Vals board, Vals MiMo scorecard).
The measuring instruments also changed. AA's September 7 revision introduced harder terminal tasks and held-out workflow automation, with private questions or answers accounting for 45% of index weight. Vals replaced major components in August and upgraded again in September. I therefore reject a linear July-to-September gap trend: changes in scores mix competitive movement with changes in the tests (AA revision, Vals version history).
The conjunction is demanding, but the boards are not independent. My calculation on twenty matched releases gives Pearson and Spearman correlations of approximately 0.93 each. The releases span June 9–September 29; their scores use the September 30 AA and September 29 Vals vintages. I use each release's highest available AA configuration and exclude ambiguous checkpoint matches. Related releases from the same laboratories and differing evaluation settings limit this selected cross-section. The correlation supports a substantial conditional chance of dual leadership, not a 93% chance of identical winners (AA input data, Vals input data).
Vals is dominated by finance and coding: normalizing its published sector weights gives 52.3% finance and 36.6% coding. Legal and tax are smaller contributors. The extra hurdle is reliable professional work, not principally a presumed Chinese disadvantage in US law. Publication timing also matters, though delay is not inevitable: MiMo's release is dated September 21 and its Vals evaluation report September 22 (Vals methodology, MiMo evaluation report).
There are credible improvement mechanisms. Under the current Vals vintage, MiMo Pro's successive generations score 33.84% and 55.20%, a 21.36-point difference. Xiaomi attributes its progress to scaled reinforcement learning. Alibaba separately announced that Qwen 4 was in training. These support further serious attempts, but neither a two-generation comparison nor a vendor roadmap establishes a durable catch-up rate (MiMo score comparison, Xiaomi technical account, Alibaba roadmap, September 22).
Compute remains a headwind, not an absolute barrier. Epoch's September 24-updated hardware analysis identifies memory and chip-performance constraints despite Huawei's accelerated roadmap. Those are modeled supplier constraints, not a measurement of all compute available to Chinese laboratories. Huawei's announced new generations and the January 13 US decision permitting conditional, case-by-case advanced-chip export reviews show why today's constraints should not be treated as fixed through the deadline (Epoch hardware analysis, Huawei announcement, US export-policy decision).
I turn this evidence into a scenario-mixture model. Each regime persists across the remaining horizon; Chinese laboratories are not treated as independent bets. The inputs below are my judgments, not measured historical frequencies. An episode means a distinct Chinese AA overall-first-place episode, including qualifying ties. The conditional success probability includes the same model also leading Vals while its public AA lead remains active.
| Competitive regime | Weight | Chinese AA-leading episodes per year | Conditional dual-lead success per episode | Future YES probability within regime |
|---|---|---|---|---|
| Persistent capability lag | 50.0% | 0.08 | 35.0% | 6% |
| Intermittent frontier parity | 35.0% | 0.38 | 55.0% | 36.8% |
| Broad Chinese breakthrough | 15% | 1.10 | 75% | 83% |
Persistent lag receives the largest weight because of the current deficits and historical moving-frontier problem. Intermittent parity receives substantial weight because of the verified near-miss and multiple credible developers. Breakthrough retains meaningful weight because large post-training, architectural, and infrastructure changes can alter the competitive order. Stronger and longer-lived leaders have a higher chance of being evaluated first on both boards.
The remaining horizon is years, calculated from the supplied timestamps. For regime , let be its weight, its episode rate, and its conditional dual-lead probability. Successful episodes have hazard . I include a judgmental competing publication-loss hazard per year:
Publication loss before success prevents a future YES; it is not being classified as NO. A documented success is absorbing and is not erased by later cessation. I also reserve a small judgmental allowance for an already-completed episode missed by incomplete archives, giving . The calculated result rounds to 28%; extra digits in the JSON record the arithmetic, not predictive certainty. The same assumptions imply a 42.1% chance of at least one future Chinese AA-leading episode before the publication-loss adjustment.
What's non-obvious
Near-parity and a persistent lag can coexist. A Chinese release can approach both leaders, then be overtaken before it leads either. The July near-miss raises the forecast, while subsequent harder evaluations prevent treating that near-miss as today's position. Neither current deficits nor generation-sized improvements support a straight-line extrapolation (Kimi's launch evaluation, AA methodology changes).
Historical evidence also needs stricter handling than a search-result summary. An archived GLM 5.3 card displayed generic “Accuracy” without a Vals Index row or overall index rank. That is insufficient to establish Vals leadership. Current model cards and launch reports can also retain different scoring vintages from the canonical board. I prioritize explicitly identified index rankings and contemporaneous archives, not unlabeled accuracy fields or dated headings alone (archived GLM card, current canonical Vals board).
Uncertainties
The largest uncertainty is future relative progress. There is no stable historical dataset from which to estimate the exact dual-leadership arrival rate. Halving my episode rates lowers the forecast to 18%; increasing them by half raises it to 35.8%. A broader 15%–45.0% assumption-sensitivity range captures changes in competitive regimes and cross-index transfer. It is not a statistical confidence interval.
Historical coverage remains incomplete. Daily, versioned leaderboard records and a timestamped evaluation queue would materially improve the estimate. Matched checkpoint and inference-setting records would also improve the cross-index comparison. The recovered archives establish particular nonqualifying states, not continuous absence (opening-period Vals archive, late-period AA archive).
Finally, training-compute access and regulatory consequences are poorly observed. Anthropic's report alleges distillation and undisclosed data routing, and reporting describes a Chinese regulatory inquiry. Those allegations are not proof that the benchmark runs used foreign models, nor proof of a capability-limiting penalty. I do not invalidate measured results or impose an extra numerical penalty without that evidence (Anthropic report, reporting on the inquiry). No prediction-market odds or public forecast aggregates were used.
Sources
- Domain Expert Search · mcp
Found 7 domain experts for 'Frontier AI capabilities Chinese models versus US models, Artificial Analysis Intelligence Index Vals Index benchmark leadership September 2026':
- Artificialanalysis · mcp
LLM Benchmarks - Artificial Analysis (as of 2026-09-30)
- Perigon · mcp
Found 350 total articles. Showing 12:
- bloomberg.com · tool
- techrepublic.com · tool
- techcrunch.com · tool
- thenextweb.com · tool
- venturebeat.com · tool
- gizmodo.com · tool
- techtimes.com · tool
- techtimes.com · tool
- techtimes.com · tool
- businessinsider.com · tool
- fortune.com · tool
- benzinga.com · tool
- sina.cn · tool
- finance.ifeng.com · tool
- sina.cn · tool
- sina.cn · tool
- 163.com · tool
- sina.cn · tool
- 163.com · tool
- stock.jrj.com.cn · tool
- Domain Expert Research Task · mcp
Job domain_expert_research_task_23d1f76978 done after 718948ms.
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- artificialanalysis.ai · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- artificialanalysis.ai · tool
- lmarena.ai · tool
- Webarchive · mcp
[ ](/web/20260819224033/https://www.vals.ai/home)
- vals.ai · tool
- AI Benchmarks for Real-World Tasks | Vals AI · openai
- AI Benchmark Leaderboards | Vals AI · openai
- vals.ai · tool
- vals.ai · tool
- vals.ai · tool
Question Details
Description
This question asks whether, at any time between the question creation date and the end of calendar year 2028 (through 2028-12-31, inclusive), a model developed by a Chinese organization will simultaneously be ranked #1 on both the Artificial Analysis Intelligence Index and the Vals Index. As of mid-2026, frontier Chinese models (including those from organizations such as DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai/GLM, MiniMax, Tencent, and Baidu) are highly competitive on public benchmarks, but the leaders of the Artificial Analysis Intelligence Index and the Vals Index are not consistently Chinese models. The gap between U.S. and Chinese frontier models has narrowed substantially in recent years. ([artificialanalysis.ai](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1))
Resolution Criteria
Resolve YES if, at any point from question creation through 2028-12-31 inclusive, the same Chinese model (or any Chinese model) is listed as the sole or joint #1 model on the public Artificial Analysis Intelligence Index leaderboard and is simultaneously listed as the sole or joint #1 model on the public Vals Index leaderboard. Resolve NO if no Chinese model satisfies both conditions simultaneously by the end of 2028. For this question, a "Chinese model" is a model whose primary developer or sponsoring organization is headquartered in the People's Republic of China (for example, DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai, MiniMax, Tencent, or Baidu). If ownership or headquarters changes, the developer's status at the time the simultaneous #1 rankings occur will be used. Primary resolution sources are the public Artificial Analysis Intelligence Index leaderboard and the public Vals Index leaderboard. If historical snapshots are needed because the leaderboards have changed by the resolution date, archived snapshots (including the Internet Archive or other reliable historical records) may be used to determine whether simultaneous leadership occurred. ([artificialanalysis.ai](https://artificialanalysis.ai/models/))
Fine Print
If either leaderboard substantially changes methodology, is renamed, or is replaced by a clear successor published by the same organization, the successor leaderboard will be used. If one leaderboard ceases publication without a clear successor before 2028-12-31, the question should be annulled unless sufficient archived evidence exists to determine whether the condition was met before publication ceased. Ties for first place count as holding the index. The Chinese model does not need to be the same model over the entire period; the question only requires that some Chinese model simultaneously holds first place on both leaderboards at least once during the resolution window.