Back to question

Forecast report

Will a Chinese model hold the AA Intelligence Index and the Vals Index simultaneously at any point between now and the end of 2028?

GeneratedAugust 31, 2026 at 2:54 PM UTC
ResolutionNot specified
Question typeBinary
Sources50

Forecast

P(Yes): 35.6%; P(No): 64.4%.

Distribution

35.6%CHANCE

Analysis

TL;DR

I estimate a 35.6% probability of YES. At the August 31, 2026 cutoff, the best Chinese models were about three points short on the Artificial Analysis index, while Kimi K3 was 9.40 percentage points short and ranked tenth on the Vals Index v2. The long window and rapid Chinese release pace create many chances, but NO remains favored because one model must lead two different boards during the same publication window.

Context

I read the resolution as requiring the same Chinese model to be sole or joint first on both boards at the same time. Different Chinese models leading the two boards would not qualify. I found no qualifying event between the July 22, 2026 opening and this forecast cutoff.

The present state is close on one board but not the other. Claude Opus 5 leads Artificial Analysis at 63.1, against 59.7 for Kimi K3 and 59.5 for GLM-5.3 in the latest detailed leaderboard snapshot. Claude Opus 5 also leads Vals v2 at 67.213%; Kimi K3 is the top Chinese model at 57.813%, ranked tenth of 48 on its official model page. (artificialanalysis.ai)

Evidence

The historical base rate is zero. Every recoverable canonical Vals state I checked—from October 2025, through May 2026, v1.2 in July 2026, and v2 in August 2026—had a non-Chinese overall leader. The closest verified case was Kimi K3 on v1.2: 74.70% and third place, just 0.445 percentage points behind Claude Fable 5 at 75.145%, with Claude Opus 5 between them. Archive coverage is not continuous, so this is strong evidence rather than proof that no brief historical lead was missed.

Artificial Analysis tells the same story. I found no verified Chinese overall winner. DeepSeek R1-0528 reached the second tier in May 2025, scoring 68 against the leading frontier around 70, while Artificial Analysis described DeepSeek as the number-two lab. Kimi K3 then reached third overall at 57 on July 17, 2026, behind Claude Fable 5 and GPT-5.6 Sol according to the official launch evaluation. (artificialanalysis.ai)

The strongest evidence for YES is the small present Artificial Analysis gap. Three Chinese developers—Moonshot, Z.ai and Alibaba—have models within five index points of first place in the late-August standings. Recent improvements are large enough to cross that gap: GLM-5.3 scores 60 versus 53 for GLM-5.2 under the same v4.1.1 methodology, a seven-point gain in roughly two months, according to Artificial Analysis. Z.ai says the models share the same base and that the improvement came from scaled post-training, though its causal account remains a vendor claim (Z.ai release note). (artificialanalysis.ai)

The strongest evidence for NO is Vals v2. Its August 13 methodology change replaced corporate-finance document analysis with a private Excel-modeling benchmark, added code migration and two legal-agent evaluations, and removed saturated SWE-bench from the coding bucket (Vals methodology). Kimi moved from a near-leader at 74.70% on v1.2 to 57.813% and tenth place on v2. That was a methodology reset, not a sudden loss of capability, but it shows that Chinese convergence was not robust across task mixes.

The boards overlap, but less than their names suggest. Artificial Analysis weights agents at 34%, coding at 24%, scientific reasoning at 24% and general capability at 18% (methodology). Vals weights finance at 54.1%, coding at 37.8% and legal work at 8.1%, based on its published GDP-weighting formula. Terminal-Bench 2.1 is the only literal component shared by both. A model can therefore reach first on Artificial Analysis through science, knowledge and broad agent performance while remaining weaker on Excel, finance or US legal work.

Release mechanics help YES. Kimi K3 was evaluated by both organizations within days of its July 16 launch, and Artificial Analysis reported that six labs had crossed 50 points by July 17 (frontier-release review). The question only needs a short overlap before a US competitor responds. But evaluation lag can also prevent overlap: a Chinese model may top one board and be displaced before the second evaluation finishes.

I combined three quantitative views:

  • A bottleneck model assigns a 68% chance that a Chinese model leads Artificial Analysis at least once, a 52% conditional chance that one such model also leads Vals during an overlapping window, and 98% leaderboard-continuity probability. This gives 0.68 × 0.52 × 0.98 = 34.65%.
  • A regime model assigns 58% to a persistent Western edge with a 7% conditional success rate, 35% to rough parity with a 73% success rate, and 7% to a clear Chinese advantage with a 94% success rate. After continuity risk, it gives 35.47%.
  • An opportunity model assumes 12 effectively distinct Chinese flagship generations or major post-training updates, each with a 3.6% dual-lead chance. It adds a 4% route through a methodology change or published tie if releases fail, then applies continuity risk, giving 37.41%.

Weighting these models 40%, 35% and 25% produces 35.63%. The inputs are judgments, not fitted parameters. The model spread is more informative than the second decimal place.

What's non-obvious

The widely repeated GLM-5.3 Vals scores of 71.48% and 69.53% are not Vals Index scores. They are simple averages across the benchmark cards then present on GLM-5.3's model page: three benchmarks in the August 19 archive, and eight on the current page. GLM-5.3 has no Vals Index entry, while the canonical v2 board explicitly lists Claude Opus 5 first and Kimi K3 as the best Chinese model. Treating GLM-5.3 as the current Vals leader materially overstates the YES probability.

A second trap is the claim that GLM-4.7 previously led Vals overall. Vals used loose wording in a social post, but its December 28, 2025 archived homepage shows GPT-5.2 first overall at 64.49% and GLM-4.7 first only among open-weight models, at 56.21% and ninth overall. The apparent historical Chinese victory disappears once the open-weight and overall boards are separated.

Uncertainties

  • Historical archives are sparse. A very short-lived ranking between captures cannot be ruled out, although I found no supporting announcement, snapshot or model card.
  • Both organizations can change methodology again. Vals v1.2 to v2 moved Kimi from a near tie to a 9.40-point deficit, so future changes could move the forecast sharply in either direction.
  • The largest unknown is whether upcoming Chinese releases improve Excel modeling, finance, legal research and long-horizon professional work together, rather than mainly coding and general reasoning.
  • Compute access remains a headwind but not a fixed ceiling. US policy shifted H200-class exports to eligible Chinese customers to case-by-case review on January 13, 2026 (Bureau of Industry and Security), and the rules could tighten or loosen again before 2028.

The data that would narrow the forecast most is a timestamped public history for both leaderboards, Vals scores for GLM-5.3 and later Chinese models on all seven v2 components, and observed evaluation turnaround times for major releases. Until those exist, 35.6% is the best balance between the real near-misses and the harder same-model, two-board requirement.

Sources

  1. Domain Expert Search · mcp

    Found 14 domain experts for 'frontier AI model benchmarks Chinese versus US models Artificial Analysis Vals Index through 2028 forecasting':

  2. Artificialanalysis · mcp

    LLM Benchmarks - Artificial Analysis (as of 2026-08-31)

  3. Webarchive · mcp

    {

  4. vals.ai · tool
  5. web.archive.org · tool
  6. web.archive.org · tool
  7. web.archive.org · tool
  8. web.archive.org · tool
  9. web.archive.org · tool
  10. web.archive.org · tool
  11. web.archive.org · tool
  12. web.archive.org · tool
  13. web.archive.org · tool
  14. web.archive.org · tool
  15. vals.ai · tool
  16. vals.ai · tool
  17. web.archive.org · tool
  18. web.archive.org · tool
  19. web.archive.org · tool
  20. web.archive.org · tool
  21. web.archive.org · tool
  22. web.archive.org · tool
  23. web.archive.org · tool
  24. web.archive.org · tool
  25. web.archive.org · tool
  26. web.archive.org · tool
  27. web.archive.org · tool
  28. web.archive.org · tool
  29. vals.ai · tool
  30. vals.ai · tool
  31. web.archive.org · tool
  32. web.archive.org · tool
  33. web.archive.org · tool
  34. vals.ai · tool
  35. vals.ai · tool
  36. web.archive.org · tool
  37. vals.ai · tool
  38. vals.ai · tool
  39. vals.ai · tool
  40. vals.ai · tool
  41. vals.ai · tool
  42. vals.ai · tool
  43. vals.ai · tool
  44. vals.ai · tool
  45. web.archive.org · tool
  46. web.archive.org · tool
  47. web.archive.org · tool
  48. web.archive.org · tool
  49. vals.ai · tool
  50. vals.ai · tool

Question Details

Description

This question asks whether, at any time between the question creation date and the end of calendar year 2028 (through 2028-12-31, inclusive), a model developed by a Chinese organization will simultaneously be ranked #1 on both the Artificial Analysis Intelligence Index and the Vals Index. As of mid-2026, frontier Chinese models (including those from organizations such as DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai/GLM, MiniMax, Tencent, and Baidu) are highly competitive on public benchmarks, but the leaders of the Artificial Analysis Intelligence Index and the Vals Index are not consistently Chinese models. The gap between U.S. and Chinese frontier models has narrowed substantially in recent years. ([artificialanalysis.ai](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1))

Resolution Criteria

Resolve YES if, at any point from question creation through 2028-12-31 inclusive, the same Chinese model (or any Chinese model) is listed as the sole or joint #1 model on the public Artificial Analysis Intelligence Index leaderboard and is simultaneously listed as the sole or joint #1 model on the public Vals Index leaderboard. Resolve NO if no Chinese model satisfies both conditions simultaneously by the end of 2028. For this question, a "Chinese model" is a model whose primary developer or sponsoring organization is headquartered in the People's Republic of China (for example, DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai, MiniMax, Tencent, or Baidu). If ownership or headquarters changes, the developer's status at the time the simultaneous #1 rankings occur will be used. Primary resolution sources are the public Artificial Analysis Intelligence Index leaderboard and the public Vals Index leaderboard. If historical snapshots are needed because the leaderboards have changed by the resolution date, archived snapshots (including the Internet Archive or other reliable historical records) may be used to determine whether simultaneous leadership occurred. ([artificialanalysis.ai](https://artificialanalysis.ai/models/))

Fine Print

If either leaderboard substantially changes methodology, is renamed, or is replaced by a clear successor published by the same organization, the successor leaderboard will be used. If one leaderboard ceases publication without a clear successor before 2028-12-31, the question should be annulled unless sufficient archived evidence exists to determine whether the condition was met before publication ceased. Ties for first place count as holding the index. The Chinese model does not need to be the same model over the entire period; the question only requires that some Chinese model simultaneously holds first place on both leaderboards at least once during the resolution window.