Back to question

Forecast report

Will a Chinese model hold the AA Intelligence Index and the Vals Index simultaneously at any point between now and the end of 2028?

GeneratedSeptember 30, 2026 at 8:54 AM UTC
ResolutionNot specified
Question typeYes / No
Sources50

Forecast

P(Yes): 28.5%; P(No): 71.5%.

Distribution

28.5%CHANCE

Analysis

TL;DR

I assign a 28% probability of YES. The strongest currently evaluated Chinese model trails the leaders on both required indices (Artificial Analysis, Vals). The opportunity is one brief overlapping lead, not permanent Chinese dominance.

Context

I use the client's September 30, 2026, 08:34 UTC cutoff and the full eligibility window, beginning July 22, 2026, at 22:46 UTC and ending December 31, 2028. I require the same identifiable Chinese model to hold first place on both overall boards during an overlapping interval. Official joint leadership counts; overlapping error bars alone do not. Ordinary differences in reasoning-effort settings do not change the underlying model's identity.

The current leaders are Claude Opus 5.5 on AA Intelligence Index v4.3.2 and Claude Sonnet 5.5 on Vals Index v2.1. Xiaomi's MiMo-V2.6-Pro is the strongest evaluated Chinese contender on both. I found no verified qualifying historical episode, but the archives do not cover every moment (AA rankings, Vals rankings, updated September 29, historical AA snapshot).

Evidence

The historical backbone favors continued Chinese catch-up without guaranteed overtaking. Epoch AI's January 2, 2026 analysis reported that every frontier model in its ECI record since 2023 had been developed in the United States. Across its 2023–late-2025 coverage, Chinese models lagged by seven months on average, with a four-to-fourteen-month range. This is an older, different index. The report does not supply an independent-trial sample size, so I use it as evidence of persistent lag, not as a fitted probability of this event (Epoch analysis).

Broader rankings give a different picture. Stanford's report describes repeated US–Chinese changes at the top of performance rankings since early 2025. Its discussion relies on broader measures, including human preference, rather than these two professional-agent indices. That rejects an extremely pessimistic prior, but does not establish leadership on the required boards (Stanford technical-performance review).

The complete set of recovered archived overall-leader observations used here is below. These are eight AA captures and six Vals index-bearing page states, spanning July 22–September 30. They are sparse, dependent observations—not fourteen independent trials. The two opening-day captures precede question creation. Scores are contemporary archived values, not current scores assigned retrospectively to older releases.

IndexCapture timestamp, UTCReported overall leaderScore
Vals2026-07-22 16:48:28 — pre-windowClaude Fable 575.14%
AA2026-07-22 17:37:08 — pre-windowClaude Fable 560 points
Vals2026-07-26 20:22:54Claude Fable 575.14%
AA2026-07-31 04:59:10Claude Opus 561 points
AA2026-08-06 07:19:46Claude Opus 561 points
Vals2026-08-15 10:01:32Claude Opus 567.21%
AA2026-08-19 02:53:23Claude Opus 563 points
AA2026-08-20 12:20:47Claude Opus 563 points
AA2026-08-21 02:48:50Claude Opus 563 points
AA2026-08-31 18:51:03Claude Opus 563 points
Vals2026-09-09 12:46:17Claude Fable 5.168.83%
Vals2026-09-21 06:55:33Claude Fable 5.168.83%
Vals2026-09-29 20:25:37Claude Sonnet 5.567.04%
AA2026-09-30 03:05:14Claude Opus 5.558 points

The strongest positive historical evidence is a near-miss. AA's July 17 report put Kimi K3 at 57, behind Fable 5 at 60. The Vals page state dated July 23, captured July 26, put Kimi at 74.70%, only 0.44 percentage points behind Fable's 75.14%. Kimi also performed strongly on private knowledge-work evaluations. Chinese proximity therefore extended beyond public mathematics and coding tests (AA's Kimi evaluation, archived Vals results).

Today's deficits are substantially larger. The September 30 AA data give Opus 5.5 maximum effort 57.6 points and MiMo Pro 46.3, an 11.3-point difference; public pages round those scores. Vals's September 29-updated board gives Sonnet 5.5 67.04% and MiMo 55.20%, an 11.84-percentage-point difference. MiMo ranks eleventh among thirty Vals entries. AA points and Vals weighted accuracy are different units and should not be averaged (AA leader, AA MiMo results, Vals board, Vals MiMo scorecard).

The measuring instruments also changed. AA's September 7 revision introduced harder terminal tasks and held-out workflow automation, with private questions or answers accounting for 45% of index weight. Vals replaced major components in August and upgraded again in September. I therefore reject a linear July-to-September gap trend: changes in scores mix competitive movement with changes in the tests (AA revision, Vals version history).

The conjunction is demanding, but the boards are not independent. My calculation on twenty matched releases gives Pearson and Spearman correlations of approximately 0.93 each. The releases span June 9–September 29; their scores use the September 30 AA and September 29 Vals vintages. I use each release's highest available AA configuration and exclude ambiguous checkpoint matches. Related releases from the same laboratories and differing evaluation settings limit this selected cross-section. The correlation supports a substantial conditional chance of dual leadership, not a 93% chance of identical winners (AA input data, Vals input data).

Vals is dominated by finance and coding: normalizing its published sector weights gives 52.3% finance and 36.6% coding. Legal and tax are smaller contributors. The extra hurdle is reliable professional work, not principally a presumed Chinese disadvantage in US law. Publication timing also matters, though delay is not inevitable: MiMo's release is dated September 21 and its Vals evaluation report September 22 (Vals methodology, MiMo evaluation report).

There are credible improvement mechanisms. Under the current Vals vintage, MiMo Pro's successive generations score 33.84% and 55.20%, a 21.36-point difference. Xiaomi attributes its progress to scaled reinforcement learning. Alibaba separately announced that Qwen 4 was in training. These support further serious attempts, but neither a two-generation comparison nor a vendor roadmap establishes a durable catch-up rate (MiMo score comparison, Xiaomi technical account, Alibaba roadmap, September 22).

Compute remains a headwind, not an absolute barrier. Epoch's September 24-updated hardware analysis identifies memory and chip-performance constraints despite Huawei's accelerated roadmap. Those are modeled supplier constraints, not a measurement of all compute available to Chinese laboratories. Huawei's announced new generations and the January 13 US decision permitting conditional, case-by-case advanced-chip export reviews show why today's constraints should not be treated as fixed through the deadline (Epoch hardware analysis, Huawei announcement, US export-policy decision).

I turn this evidence into a scenario-mixture model. Each regime persists across the remaining horizon; Chinese laboratories are not treated as independent bets. The inputs below are my judgments, not measured historical frequencies. An episode means a distinct Chinese AA overall-first-place episode, including qualifying ties. The conditional success probability includes the same model also leading Vals while its public AA lead remains active.

Competitive regimeWeightChinese AA-leading episodes per yearConditional dual-lead success per episodeFuture YES probability within regime
Persistent capability lag50.0%0.0835.0%6%
Intermittent frontier parity35.0%0.3855.0%36.8%
Broad Chinese breakthrough15%1.1075%83%

Persistent lag receives the largest weight because of the current deficits and historical moving-frontier problem. Intermittent parity receives substantial weight because of the verified near-miss and multiple credible developers. Breakthrough retains meaningful weight because large post-training, architectural, and infrastructure changes can alter the competitive order. Stronger and longer-lived leaders have a higher chance of being evaluated first on both boards.

The remaining horizon is T=2.255T=2.255 years, calculated from the supplied timestamps. For regime ss, let wsw_s be its weight, λs\lambda_s its episode rate, and qsq_s its conditional dual-lead probability. Successful episodes have hazard hs=λsqsh_s=\lambda_s q_s. I include a judgmental competing publication-loss hazard δ=0.02\delta=0.02 per year:

Pfuture=∑swshshs+δ(1−e−(hs+δ)T).P_{\mathrm{future}}=\sum_s w_s\frac{h_s}{h_s+\delta}\left(1-e^{-(h_s+\delta)T}\right).

Publication loss before success prevents a future YES; it is not being classified as NO. A documented success is absorbing and is not erased by later cessation. I also reserve a small judgmental allowance r=0.002r=0.002 for an already-completed episode missed by incomplete archives, giving P(YES)=r+(1−r)PfutureP(\mathrm{YES})=r+(1-r)P_{\mathrm{future}}. The calculated result rounds to 28%; extra digits in the JSON record the arithmetic, not predictive certainty. The same assumptions imply a 42.1% chance of at least one future Chinese AA-leading episode before the publication-loss adjustment.

What's non-obvious

Near-parity and a persistent lag can coexist. A Chinese release can approach both leaders, then be overtaken before it leads either. The July near-miss raises the forecast, while subsequent harder evaluations prevent treating that near-miss as today's position. Neither current deficits nor generation-sized improvements support a straight-line extrapolation (Kimi's launch evaluation, AA methodology changes).

Historical evidence also needs stricter handling than a search-result summary. An archived GLM 5.3 card displayed generic “Accuracy” without a Vals Index row or overall index rank. That is insufficient to establish Vals leadership. Current model cards and launch reports can also retain different scoring vintages from the canonical board. I prioritize explicitly identified index rankings and contemporaneous archives, not unlabeled accuracy fields or dated headings alone (archived GLM card, current canonical Vals board).

Uncertainties

The largest uncertainty is future relative progress. There is no stable historical dataset from which to estimate the exact dual-leadership arrival rate. Halving my episode rates lowers the forecast to 18%; increasing them by half raises it to 35.8%. A broader 15%–45.0% assumption-sensitivity range captures changes in competitive regimes and cross-index transfer. It is not a statistical confidence interval.

Historical coverage remains incomplete. Daily, versioned leaderboard records and a timestamped evaluation queue would materially improve the estimate. Matched checkpoint and inference-setting records would also improve the cross-index comparison. The recovered archives establish particular nonqualifying states, not continuous absence (opening-period Vals archive, late-period AA archive).

Finally, training-compute access and regulatory consequences are poorly observed. Anthropic's report alleges distillation and undisclosed data routing, and reporting describes a Chinese regulatory inquiry. Those allegations are not proof that the benchmark runs used foreign models, nor proof of a capability-limiting penalty. I do not invalidate measured results or impose an extra numerical penalty without that evidence (Anthropic report, reporting on the inquiry). No prediction-market odds or public forecast aggregates were used.

Sources

  1. Domain Expert Search · mcp

    Found 7 domain experts for 'Frontier AI capabilities Chinese models versus US models, Artificial Analysis Intelligence Index Vals Index benchmark leadership September 2026':

  2. Artificialanalysis · mcp

    LLM Benchmarks - Artificial Analysis (as of 2026-09-30)

  3. Perigon · mcp

    Found 350 total articles. Showing 12:

  4. bloomberg.com · tool
  5. techrepublic.com · tool
  6. techcrunch.com · tool
  7. thenextweb.com · tool
  8. venturebeat.com · tool
  9. gizmodo.com · tool
  10. techtimes.com · tool
  11. techtimes.com · tool
  12. techtimes.com · tool
  13. businessinsider.com · tool
  14. fortune.com · tool
  15. benzinga.com · tool
  16. sina.cn · tool
  17. finance.ifeng.com · tool
  18. sina.cn · tool
  19. sina.cn · tool
  20. 163.com · tool
  21. sina.cn · tool
  22. 163.com · tool
  23. stock.jrj.com.cn · tool
  24. Domain Expert Research Task · mcp

    Job domain_expert_research_task_23d1f76978 done after 718948ms.

  25. web.archive.org · tool
  26. web.archive.org · tool
  27. web.archive.org · tool
  28. artificialanalysis.ai · tool
  29. web.archive.org · tool
  30. web.archive.org · tool
  31. web.archive.org · tool
  32. web.archive.org · tool
  33. web.archive.org · tool
  34. web.archive.org · tool
  35. web.archive.org · tool
  36. web.archive.org · tool
  37. web.archive.org · tool
  38. web.archive.org · tool
  39. web.archive.org · tool
  40. web.archive.org · tool
  41. web.archive.org · tool
  42. artificialanalysis.ai · tool
  43. lmarena.ai · tool
  44. Webarchive · mcp

    [ ](/web/20260819224033/https://www.vals.ai/home)

  45. vals.ai · tool
  46. AI Benchmarks for Real-World Tasks | Vals AI · openai
  47. AI Benchmark Leaderboards | Vals AI · openai
  48. vals.ai · tool
  49. vals.ai · tool
  50. vals.ai · tool

Question Details

Description

This question asks whether, at any time between the question creation date and the end of calendar year 2028 (through 2028-12-31, inclusive), a model developed by a Chinese organization will simultaneously be ranked #1 on both the Artificial Analysis Intelligence Index and the Vals Index. As of mid-2026, frontier Chinese models (including those from organizations such as DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai/GLM, MiniMax, Tencent, and Baidu) are highly competitive on public benchmarks, but the leaders of the Artificial Analysis Intelligence Index and the Vals Index are not consistently Chinese models. The gap between U.S. and Chinese frontier models has narrowed substantially in recent years. ([artificialanalysis.ai](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1))

Resolution Criteria

Resolve YES if, at any point from question creation through 2028-12-31 inclusive, the same Chinese model (or any Chinese model) is listed as the sole or joint #1 model on the public Artificial Analysis Intelligence Index leaderboard and is simultaneously listed as the sole or joint #1 model on the public Vals Index leaderboard. Resolve NO if no Chinese model satisfies both conditions simultaneously by the end of 2028. For this question, a "Chinese model" is a model whose primary developer or sponsoring organization is headquartered in the People's Republic of China (for example, DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai, MiniMax, Tencent, or Baidu). If ownership or headquarters changes, the developer's status at the time the simultaneous #1 rankings occur will be used. Primary resolution sources are the public Artificial Analysis Intelligence Index leaderboard and the public Vals Index leaderboard. If historical snapshots are needed because the leaderboards have changed by the resolution date, archived snapshots (including the Internet Archive or other reliable historical records) may be used to determine whether simultaneous leadership occurred. ([artificialanalysis.ai](https://artificialanalysis.ai/models/))

Fine Print

If either leaderboard substantially changes methodology, is renamed, or is replaced by a clear successor published by the same organization, the successor leaderboard will be used. If one leaderboard ceases publication without a clear successor before 2028-12-31, the question should be annulled unless sufficient archived evidence exists to determine whether the condition was met before publication ceased. Ties for first place count as holding the index. The Chinese model does not need to be the same model over the entire period; the question only requires that some Chinese model simultaneously holds first place on both leaderboards at least once during the resolution window.