Forecast report
Will a Chinese model hold the AA Intelligence Index and the Vals Index simultaneously at any point between now and the end of 2028?
Forecast
P(Yes): 41.1%; P(No): 58.9%.
Distribution
Analysis
TL;DR
I estimate a 41.1% chance of YES. China is close on the Artificial Analysis leaderboard, but Vals v2 is the bottleneck: Kimi K3 trails by 9.40 percentage points after sitting just 0.44 points behind under v1.2 in July. The long window and several strong Chinese labs create many shots, but the verified Vals history has no Chinese overall leader and U.S. labs keep moving the target.
Context
The condition has not occurred. The August 31 Artificial Analysis data has Claude Opus 5 first at 63.1 index points, versus 59.7 for Kimi K3 and 59.5 for GLM-5.3. The Vals Index updated August 19 has Claude Opus 5 first at 67.21%, Claude Fable 5 second at 66.04%, and GPT-5.6 Sol third at 63.71%; Kimi K3 is the best confirmed Chinese model at 57.81%, rank 10 of 48. These are different units and should not be compared directly.
The AA Intelligence Index v4.1.1 weights agents at 34%, coding at 24%, scientific reasoning at 24%, and general capability at 18%. Vals v2 combines seven benchmarks across finance, coding, and law, five of them private, with finance receiving 8.0 of 14.8 weight units. One particular Chinese model must be listed first on both boards at the same time. Ties count.
Evidence
The cleanest historical backbone is Vals. These are all the archived checkpoints I could verify with both the overall leader and the best Chinese score. Scores and gaps are percentage points; rows from different methodologies are not directly comparable.
| Published checkpoint | Version | Overall leader | Best Chinese model | Gap |
|---|---|---|---|---|
| 2025-10-09 | Unlabeled | Claude Sonnet 4.5, 66.60 | GLM 4.5, 42.00 | 24.60 |
| 2025-10-29 | Unlabeled | Claude Sonnet 4.5, 66.70 | GLM 4.6, 47.20 | 19.50 |
| 2025-12-12 | Unlabeled | GPT-5.2, 64.49 | DeepSeek V3.2, 49.39 | 15.10 |
| 2026-01-12 | Unlabeled | GPT-5.2, 64.49 | GLM 4.7, 56.21 | 8.28 |
| 2026-02-10 | Unlabeled | Claude Opus 4.6, 65.98 | Kimi K2.5, 59.74 | 6.24 |
| 2026-03-20 | Unlabeled | Claude Sonnet 4.6, 67.74 | GLM 5, 61.41 | 6.33 |
| 2026-04-21 | Unlabeled | Claude Opus 4.7, 71.47 | Kimi K2.6, 63.94 | 7.53 |
| 2026-05-13 | v1.1 | GPT-5.5, 67.62 | DeepSeek V4, 56.23 | 11.39 |
| 2026-06-10 | v1.2 | Claude Fable 5, 75.14 | MiniMax-M3, 58.94 | 16.20 |
| 2026-07-21 | v1.2 | Claude Fable 5, 75.14 | Kimi K3, 74.70 | 0.44 |
| 2026-08-18 | v2 | Claude Opus 5, 67.21 | Kimi K3, 57.81 | 9.40 |
Zero of these 11 checkpoints had a Chinese leader. Six had a gap below ten points, and one was a genuine near miss. The sequence is not a smooth catch-up curve. The v2 methodology reset alone took Kimi from 0.44 points behind to 9.40 points behind.
Artificial Analysis shows the same mix of rapid progress and a moving target. Its July 17 assessment put Kimi K3 third at 57, behind Fable 5 at 60 and GPT-5.6 Sol at 59. Under the current common methodology, Kimi K2.6 to Kimi K3 rose from 45 to 60, GLM-5.2 to GLM-5.3 rose from 53 to 60, and Qwen3.7 Max to Qwen3.8 Max rose from 47 to 58. These are selected successful generations, not a trend that can be extrapolated indefinitely. The AA leader also changed from Fable 5 to Opus 5 between July 22 and July 24.
Vals is harder because the weakness is uneven. Kimi K3’s current component ranks include third on Vibe Code Bench and Terminal-Bench 2.1, but seventh on Excel modeling, tenth on Finance Agent, sixth on Legal Research, seventh on Harvey’s Legal Agent Benchmark, and thirtieth on Code Migration. Yet Kimi is fifth on Scale’s private, 1,000-task MCP Atlas tool-use benchmark, at 82.3%. I read this as a professional reliability and deliverable-quality gap, not a simple lack of coding or tool-use ability.
Compute remains a drag, not a hard ceiling. Huawei’s own roadmap targets the training-oriented Ascend 950DT for the fourth quarter of 2026 and Ascend 960 for the fourth quarter of 2027, while U.S. reporting shows both continuing chip restrictions and gaps involving remote access to advanced systems (Huawei, CNBC). Vendor roadmaps can slip, but partial execution would give Chinese labs more training capacity before resolution.
I combined three quantitative views. First, a Jeffreys-prior monthly hazard model applied to zero Chinese Vals leaders in the 11 verified checkpoints gives 46.5% for at least one leading month during the next 28 months. I raise the underlying Vals-lead estimate to 55.0% because of the July near miss and the number of competitive laboratories, then apply 78% for the same model also leading AA, 90% for overlapping publication windows, and 98% for both boards remaining usable. That gives 37.8%.
Second, I model 15 effective Chinese flagship waves through 2028, fewer than the raw release count because releases share talent, compute, and techniques. Joint-success hazards of 2% for the remaining 2026 waves, 3% for 2027, and 5% for 2028 produce 44.1%. Third, a regime model assigns 20% to repeated Chinese frontier advantage with an 85% conditional YES rate, 50% to rough parity with a 45.0% YES rate, 25% to a persistent U.S. professional-capability lead with a 10% YES rate, and 5% to publication disruption with a 5% YES rate. That gives 42.3%. Weighting the three models 40:30:30 gives 41.1%.
What's non-obvious
The indices are not independent. Treating them as two separate coin flips is too bearish because their frontier rankings overlap and both reward agentic coding and tool use. But treating the July near miss as the current baseline is too bullish. The v2 reset is load-bearing: it replaced saturated and narrower tasks with spreadsheet modeling, code migration, legal research, and long-horizon work products.
The race is also max versus max. China has several credible laboratories, but it must beat the best release from Anthropic, OpenAI, Google, Meta, and xAI. This raises the value of China’s many shots while also making raw release counts misleading. The relevant number is effective independent release waves, not every checkpoint or minor update.
Uncertainties
- Vals v2 was only 18 days old at the forecast cutoff. It has one mature public vintage, so its 9.40-point Chinese gap may not be persistent.
- GLM-5.3, which nearly matches Kimi on AA, did not yet have a published full Vals Index rank. Its result could materially change the near-term view.
- Artificial Analysis repeatedly regrades older models when methodology changes. A complete daily archive of displayed rankings, rather than current-methodology backcasts, would produce a cleaner base rate.
- Future index revisions, evaluation lags, chip-access policy, and possible leaderboard discontinuation create a wide outcome range. Full timestamped score histories and several more Vals v2 evaluations of new Chinese flagships would close most of the current information gap.
Sources
- Domain Expert Search · mcp
Found 14 domain experts for 'frontier AI models China United States benchmark leaderboards Artificial Analysis Vals model development trends compute and release cadence':
- Artificialanalysis · mcp
LLM Benchmarks - Artificial Analysis (as of 2026-08-31)
- Scale Seal · mcp
SEAL leaderboards (26):
- Epoch · mcp
Epoch 'frontier' models: 2 shown (newest first), published 2024-01-17 → 2024-08-29.
- bigmodel.cn · tool
- epoch.ai · tool
- huggingface.co · tool
- huggingface.co · tool
- github.com · tool
- developer.download.nvidia.com · tool
- cloud.google.com · tool
- Webarchive · mcp
{
- vals.ai · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- Benchmarks · openai
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
- web.archive.org · tool
Question Details
Description
This question asks whether, at any time between the question creation date and the end of calendar year 2028 (through 2028-12-31, inclusive), a model developed by a Chinese organization will simultaneously be ranked #1 on both the Artificial Analysis Intelligence Index and the Vals Index. As of mid-2026, frontier Chinese models (including those from organizations such as DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai/GLM, MiniMax, Tencent, and Baidu) are highly competitive on public benchmarks, but the leaders of the Artificial Analysis Intelligence Index and the Vals Index are not consistently Chinese models. The gap between U.S. and Chinese frontier models has narrowed substantially in recent years. ([artificialanalysis.ai](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1))
Resolution Criteria
Resolve YES if, at any point from question creation through 2028-12-31 inclusive, the same Chinese model (or any Chinese model) is listed as the sole or joint #1 model on the public Artificial Analysis Intelligence Index leaderboard and is simultaneously listed as the sole or joint #1 model on the public Vals Index leaderboard. Resolve NO if no Chinese model satisfies both conditions simultaneously by the end of 2028. For this question, a "Chinese model" is a model whose primary developer or sponsoring organization is headquartered in the People's Republic of China (for example, DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai, MiniMax, Tencent, or Baidu). If ownership or headquarters changes, the developer's status at the time the simultaneous #1 rankings occur will be used. Primary resolution sources are the public Artificial Analysis Intelligence Index leaderboard and the public Vals Index leaderboard. If historical snapshots are needed because the leaderboards have changed by the resolution date, archived snapshots (including the Internet Archive or other reliable historical records) may be used to determine whether simultaneous leadership occurred. ([artificialanalysis.ai](https://artificialanalysis.ai/models/))
Fine Print
If either leaderboard substantially changes methodology, is renamed, or is replaced by a clear successor published by the same organization, the successor leaderboard will be used. If one leaderboard ceases publication without a clear successor before 2028-12-31, the question should be annulled unless sufficient archived evidence exists to determine whether the condition was met before publication ceased. Ties for first place count as holding the index. The Chinese model does not need to be the same model over the entire period; the question only requires that some Chinese model simultaneously holds first place on both leaderboards at least once during the resolution window.