# Will a Chinese model hold the AA Intelligence Index and the Vals Index simultaneously at any point between now and the end of 2028?

Canonical URL: https://preseen.com/reports/822e5c27-8a29-40a6-9adc-bf640dc31b6c/will-a-chinese-model-hold-the-aa-intelligence-index-and-the-vals-index-simultane
Markdown URL: https://preseen.com/reports/822e5c27-8a29-40a6-9adc-bf640dc31b6c/markdown

## Forecast

P(Yes): 53.5%; P(No): 46.5%.

Generated: August 1, 2026 at 4:43 PM UTC
Forecast model: gpt-5.6-sol
Research model: gpt-5.6-sol

## Analysis

## TL;DR

I forecast a **53.5%** chance of YES. Kimi K3 is only 0.44 percentage points behind first on Vals but 3.6 index points behind first on Artificial Analysis ([Vals, updated July 31](https://www.vals.ai/benchmarks); [Artificial Analysis leaderboard](https://artificialanalysis.ai/leaderboards/models)). **Artificial Analysis is the gate**: if a Chinese model reaches first there, simultaneous Vals leadership becomes likely.

## Context

I use the stricter reading that one model must be listed first on both boards at the same time. The event had not occurred by August 1, 2026: when the window opened, Kimi K3 was second on Vals and third on Artificial Analysis, behind the leaders on both boards ([Vals’s Kimi evaluation](https://www.linkedin.com/posts/vals-ai_kimi-k3-is-now-the-2-overall-model-on-the-activity-7483601379094130688-VGDq); [Artificial Analysis’s July 17 evaluation](https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5/)).

The current Artificial Analysis leader is Claude Opus 5 at 60.7 index points; Kimi K3, the top Chinese model, has 57.1. Vals lists Claude Fable 5 at 75.14%, Claude Opus 5 at 74.82%, and Kimi K3 at 74.70%, among 41 tested models as of July 31 ([Artificial Analysis leaderboard](https://artificialanalysis.ai/leaderboards/models); [Vals leaderboard](https://www.vals.ai/benchmarks)).

## Evidence

The exact historical base rate is zero, but the record is short and unstable. The complete set of rank-visible monthly checkpoints I could verify from March through July 2026 contains five observations per board, covering March 10 through July 31; every overall leader was American. This is a checkpoint sample, not a continuous transition log.

| Checkpoint | Artificial Analysis leader and best Chinese model | Vals leader |
|---|---|---|
| March 2026 | Gemini 3.1 Pro Preview and GPT-5.4 tied at 57 points; GLM-5 had 50 ([March 25 archive](https://web.archive.org/web/20260325142450/https://artificialanalysis.ai/models)) | Claude Sonnet 4.6 ([March 10 archive](https://web.archive.org/web/20260310142050/https://www.vals.ai/benchmarks/vals_index)) |
| April 2026 | GPT-5.5 led at 60; Kimi K2.6 and MiMo-V2.5-Pro had 54 ([April 30 archive](https://web.archive.org/web/20260430033828/https://artificialanalysis.ai/models)) | Claude Sonnet 4.6 ([April 15 archive](https://web.archive.org/web/20260415103942/https://www.vals.ai/benchmarks/vals_index)) |
| May 2026 | Claude Opus 4.8 led at 61; Kimi K2.6 and MiMo-V2.5-Pro had 54 ([May 30 archive](https://web.archive.org/web/20260530090232/https://artificialanalysis.ai/models)) | GPT-5.5 ([May 23 archive](https://web.archive.org/web/20260523132840/https://www.vals.ai/benchmarks/vals_index)) |
| June 2026 | Claude Fable 5 led at 60; GLM-5.2 had 51 under v4.1 ([June 30 archive](https://web.archive.org/web/20260630050414/https://artificialanalysis.ai/models)) | Claude Fable 5 at 75.15% ([June 11 archive](https://web.archive.org/web/20260611131529/https://www.vals.ai/benchmarks/vals_index)) |
| July 2026 | Claude Opus 5 led at 61; Kimi K3 had 57 ([July 31 archive](https://web.archive.org/web/20260731195920/https://artificialanalysis.ai/models)) | Claude Fable 5 at 75.14%; Kimi K3 was third at 74.70% ([July 27 archive](https://web.archive.org/web/20260727143418/https://www.vals.ai/benchmarks/vals_index)) |

The present gap is still small enough for one release to cross. Kimi K3 gained about 12.9 Artificial Analysis points over Kimi K2.6 while using 21% fewer output tokens; on Vals it gained roughly 20 percentage points in less than three months ([Artificial Analysis’s Kimi analysis](https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5/); [Vals’s Kimi evaluation](https://www.linkedin.com/posts/vals-ai_kimi-k3-is-now-the-2-overall-model-on-the-activity-7483601379094130688-VGDq)). These jumps cannot be projected linearly, but they show that 3.6 Artificial Analysis points is a normal-generation-sized hurdle, not a technological chasm.

The boards are also **not independent lotteries**. I matched the best Artificial Analysis configuration for 35 model families against the July Vals table. The August 1 cross-section produced Pearson correlation of 0.955 and Spearman rank correlation of 0.941; the top four families on both were Opus 5, Fable 5, GPT-5.6 Sol, and Kimi K3, in different orders ([Artificial Analysis data](https://artificialanalysis.ai/leaderboards/models); [Vals score table](https://benchlm.ai/benchmarks/valsindex)). This is a selected current cross-section, not a causal result, but it supports an 84% conditional estimate that a Chinese Artificial Analysis leader would also lead Vals during an overlapping evaluation window.

The broader evidence is genuinely mixed. Stanford’s 2026 AI Index says the top U.S. model led the top Chinese model by only 2.7% on its Arena-based measure in March 2026, after the countries had traded places at the top ([Stanford technical-performance chapter](https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance)). Epoch AI instead finds that Chinese models lagged the U.S. frontier by an average of seven months from 2023 through its January 2026 publication date ([Epoch AI](https://epoch.ai/data-insights/us-vs-china-eci)). NIST’s May 1 evaluation placed DeepSeek V4 about eight months behind the U.S. frontier using 16 benchmarks across 35 models, including held-out software-engineering and reasoning tests ([NIST CAISI](https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro)). I read the disagreement as benchmark dependence: China is near parity on coding and agentic work but remains less consistently competitive on broad held-out suites.

The moving target prevents a high-confidence YES. Artificial Analysis recorded four frontier launches in eight days in July, and Claude Opus 5 displaced Kimi only eight days after Kimi’s release ([Artificial Analysis release review](https://artificialanalysis.ai/articles/four-frontier-launches-in-eight-days-six-labs-now-field-a-model-above-50-on-the-artificial-analysis-intelligence-index)). Anthropic, OpenAI, Google, Meta, and xAI will keep generating new ceilings while Moonshot, Alibaba, DeepSeek, Z.ai, MiniMax, Tencent, and others try to cross them.

There are credible extra shots, but they are weak signals rather than scored evidence. Alibaba previewed the 2.4-trillion-parameter Qwen3.8-Max and claimed it ranked second only to Fable 5 without publishing an independent benchmark table; Reuters reported that MiniMax was developing a 2.7-trillion-parameter model for possible release in the third quarter of 2026 ([Alibaba’s July 20 announcement](https://www.alibabagroup.com/en-US/document-2016703577908576256); [Reuters report](https://www.marketscreener.com/news/china-s-minimax-plans-to-launch-giant-2-7-trillion-parameter-model-ce7f5ed9d98ffe27)). Parameter count alone carries little forecast weight.

Compute and access remain modest downward pressures. The United States moved H200-class exports to approved Chinese customers to case-by-case licensing on January 13, 2026, easing but not removing the hardware constraint ([U.S. Bureau of Industry and Security](https://www.bis.gov/press-release/department-commerce-revises-license-review-policy-semiconductors-exported-china)). Reuters also reported July 7 discussions in Beijing about restricting overseas access to the most advanced Chinese models; no final policy was reported ([Reuters report](https://www.marketscreener.com/news/beijing-is-looking-at-curbing-overseas-access-to-china-s-top-ai-models-sources-say-ce7f5edbd08bf321)). Either restriction could slow evaluation, though open weights and domestic APIs reduce the risk.

My quantitative synthesis uses three models. First, a gating model assigns a 69.0% chance that some Chinese model reaches first on Artificial Analysis, an 84% chance of simultaneous Vals leadership conditional on that, a 95% evaluation-overlap factor, and a 97% leaderboard-continuity factor: `0.69 × 0.84 × 0.95 × 0.97 = 0.534`. Second, a scenario model assigns 35.0% weight to a persistent U.S. lead with a 25% event chance, 50.0% to rough parity with a 65.0% event chance, and 15% to a Chinese breakthrough with a 90% event chance, producing 54.8%. Third, conditional hazards of 18% for the remainder of 2026, 25% for 2027, and 24% for 2028 produce `1 − 0.82 × 0.75 × 0.76 = 53.3%`. The approaches cluster near 54%; I reduce the combined estimate slightly for access, archival, and annulment risk, yielding 53.5%.

## What's non-obvious

Kimi’s raw Artificial Analysis rank of seventh overstates the distance. Several entries above it are different reasoning-effort settings of the same Opus 5 and GPT-5.6 Sol families. Collapsing configurations, Kimi is fourth among distinct model families and needs to beat one 60.7-point ceiling, not six independent competitors ([Artificial Analysis leaderboard](https://artificialanalysis.ai/leaderboards/models)).

The historical Vals record also contains a misleading signal. One Vals post said GLM-5.2 “ranks #1 on the Vals Index,” but an adjacent post from the same release described it as the number-one open-weight model and fifth across all models ([full-results post](https://www.linkedin.com/posts/vals-ai_full-results-for-glm-52-are-here-this-activity-7473441362927874048-RItq); [open-weight clarification](https://www.linkedin.com/posts/vals-ai_glm-52-is-the-new-open-weight-sota-on-the-activity-7473104327474192385-F8VH)). Its published 65.02% score was below the contemporary closed-model leaders, so I do not count this as prior Chinese overall leadership.

## Uncertainties

- Historical coverage is incomplete. Older Vals pages were client-rendered, and monthly archives could miss a leadership spell lasting only hours or days. I found no such spell, but silence is not proof that none occurred.
- Both targets can change beneath the forecast. Artificial Analysis introduced v4.1 on June 15, while Vals removed law, added Vibe Code Bench, switched to Finance Agent v2, and upgraded Terminal-Bench during May 2026 ([Artificial Analysis v4.1](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1/); [Vals methodology archive](https://web.archive.org/web/20260523132840/https://www.vals.ai/benchmarks/vals_index)). Another major revision could move the probability by more than ten percentage points.
- The strongest near-term Chinese candidates remain unscored or partly scored. Independent Artificial Analysis and Vals results for Qwen3.8-Max, the reported MiniMax model, and the next Kimi generation would narrow the forecast most.
- The resolution language has a small ambiguity over whether different Chinese models could lead the two boards. I use the stricter same-model interpretation; allowing different models would raise the estimate.

My subjective 80% interval for the true forecast probability is 35.0%–72%. The width reflects unknown future model generations and benchmark revisions, not uncertainty about the current rankings.

## Sources

- Domain Expert Search (mcp)
  > Found 14 domain experts for 'China versus US frontier AI models, benchmark leaderboards, model release trajectories, and probability of Chinese model leadership by 2028':
- Artificialanalysis (mcp)
  > Tool artificialanalysis_get_llm_benchmarks on artificialanalysis returned an error:
- [errors.pydantic.dev](https://errors.pydantic.dev/2.13/v/literal_error) (tool)
- Epoch (mcp)
  > Epoch 'frontier' models: 2 shown (newest first), published 2024-01-17 → 2024-08-29.
- [bigmodel.cn](https://bigmodel.cn/) (tool)
- [epoch.ai](https://epoch.ai/data/frontier_ai_models.csv) (tool)
- [huggingface.co](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1) (tool)
- [huggingface.co](https://huggingface.co/meta-llama/Meta-Llama-3.1-8B/blob/main/LICENSE) (tool)
- [github.com](https://github.com/meta-llama/llama-recipes/blob/main/src/lla…) (tool)
- [developer.download.nvidia.com](https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf) (tool)
- [cloud.google.com](https://cloud.google.com/vertex-ai/generative-ai/docs/learn/models) (tool)
- Webarchive (mcp)
  > {
- [vals.ai](https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260727143418/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260611131529/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260523132840/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260415103942/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260310142050/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20260101183231/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20251215085159/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20251116035433/https://www.vals.ai/benchmarks/vals_index) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://www.vals.ai/benchmarks/vals_index) (tool)
- [artificialanalysis.ai](https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260731195920/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260630050414/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260530090232/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260430033828/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260325142450/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260228174954/https://artificialanalysis.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20260131135238/https://artificialanalysis.ai/models) (tool)
- [vals.ai](https://www.vals.ai/home) (tool)
- [vals.ai](https://www.vals.ai/benchmarks) (tool)
- [vals.ai](https://www.vals.ai/models/anthropic_claude-haiku-4-5-20251001-thinking) (tool)
- [vals.ai](https://www.vals.ai/application-reports/alexi) (tool)
- [vals.ai](https://www.vals.ai/comparison) (tool)
- [vals.ai](https://www.vals.ai/methodology) (tool)
- [vals.ai](https://www.vals.ai/product) (tool)
- [vals.ai](https://www.vals.ai/about) (tool)
- [vals.ai](https://www.vals.ai/benchmarks/corp_fin_v2) (tool)
- [vals.ai](https://www.vals.ai/benchmarks/finance_agent) (tool)
- [vals.ai](https://www.vals.ai/benchmarks/case_law_v2) (tool)
- [vals.ai](https://www.vals.ai/benchmarks/swebench) (tool)
- [vals.ai](https://www.vals.ai/benchmarks/terminal-bench) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://www.platform.vals.ai/privacy/Product_Privacy_Statement.pdf) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://www.vals.ai) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://twitter.com/_valsai) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://www.linkedin.com/company/vals-ai/posts?feedView=all) (tool)
- [vals.ai](https://www.vals.ai/models) (tool)
- [web.archive.org](https://web.archive.org/web/20251027191324/https://www.vals.ai/product) (tool)
- [vals.ai](https://www.vals.ai/updates) (tool)

## Question Details

This question asks whether, at any time between the question creation date and the end of calendar year 2028 (through 2028-12-31, inclusive), a model developed by a Chinese organization will simultaneously be ranked #1 on both the Artificial Analysis Intelligence Index and the Vals Index. As of mid-2026, frontier Chinese models (including those from organizations such as DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai/GLM, MiniMax, Tencent, and Baidu) are highly competitive on public benchmarks, but the leaders of the Artificial Analysis Intelligence Index and the Vals Index are not consistently Chinese models. The gap between U.S. and Chinese frontier models has narrowed substantially in recent years. (artificialanalysis.ai)

### Resolution Criteria

Resolve YES if, at any point from question creation through 2028-12-31 inclusive, the same Chinese model (or any Chinese model) is listed as the sole or joint #1 model on the public Artificial Analysis Intelligence Index leaderboard and is simultaneously listed as the sole or joint #1 model on the public Vals Index leaderboard. Resolve NO if no Chinese model satisfies both conditions simultaneously by the end of 2028. For this question, a "Chinese model" is a model whose primary developer or sponsoring organization is headquartered in the People's Republic of China (for example, DeepSeek, Alibaba/Qwen, Moonshot AI, Z.ai, MiniMax, Tencent, or Baidu). If ownership or headquarters changes, the developer's status at the time the simultaneous #1 rankings occur will be used. Primary resolution sources are the public Artificial Analysis Intelligence Index leaderboard and the public Vals Index leaderboard. If historical snapshots are needed because the leaderboards have changed by the resolution date, archived snapshots (including the Internet Archive or other reliable historical records) may be used to determine whether simultaneous leadership occurred. (artificialanalysis.ai)

### Fine Print

If either leaderboard substantially changes methodology, is renamed, or is replaced by a clear successor published by the same organization, the successor leaderboard will be used. If one leaderboard ceases publication without a clear successor before 2028-12-31, the question should be annulled unless sufficient archived evidence exists to determine whether the condition was met before publication ceased. Ties for first place count as holding the index. The Chinese model does not need to be the same model over the entire period; the question only requires that some Chinese model simultaneously holds first place on both leaderboards at least once during the resolution window.
