Loading questions
Loading questions
Generated Aug 31, 2026, 2:54 PM
I estimate a 35.6% probability of YES. At the August 31, 2026 cutoff, the best Chinese models were about three points short on the Artificial Analysis index, while Kimi K3 was 9.40 percentage points short and ranked tenth on the Vals Index v2. The long window and rapid Chinese release pace create many chances, but NO remains favored because one model must lead two different boards during the same publication window.
I read the resolution as requiring the same Chinese model to be sole or joint first on both boards at the same time. Different Chinese models leading the two boards would not qualify. I found no qualifying event between the July 22, 2026 opening and this forecast cutoff.
The present state is close on one board but not the other. Claude Opus 5 leads Artificial Analysis at 63.1, against 59.7 for Kimi K3 and 59.5 for GLM-5.3 in the latest detailed leaderboard snapshot. Claude Opus 5 also leads Vals v2 at 67.213%; Kimi K3 is the top Chinese model at 57.813%, ranked tenth of 48 on its official model page. (artificialanalysis.ai)
The historical base rate is zero. Every recoverable canonical Vals state I checked—from October 2025, through May 2026, v1.2 in July 2026, and v2 in August 2026—had a non-Chinese overall leader. The closest verified case was Kimi K3 on v1.2: 74.70% and third place, just 0.445 percentage points behind Claude Fable 5 at 75.145%, with Claude Opus 5 between them. Archive coverage is not continuous, so this is strong evidence rather than proof that no brief historical lead was missed.
Artificial Analysis tells the same story. I found no verified Chinese overall winner. DeepSeek R1-0528 reached the second tier in May 2025, scoring 68 against the leading frontier around 70, while Artificial Analysis described DeepSeek as the number-two lab. Kimi K3 then reached third overall at 57 on July 17, 2026, behind Claude Fable 5 and GPT-5.6 Sol according to the official launch evaluation. (artificialanalysis.ai)
The strongest evidence for YES is the small present Artificial Analysis gap. Three Chinese developers—Moonshot, Z.ai and Alibaba—have models within five index points of first place in the late-August standings. Recent improvements are large enough to cross that gap: GLM-5.3 scores 60 versus 53 for GLM-5.2 under the same v4.1.1 methodology, a seven-point gain in roughly two months, according to Artificial Analysis. Z.ai says the models share the same base and that the improvement came from scaled post-training, though its causal account remains a vendor claim (Z.ai release note). (artificialanalysis.ai)
The strongest evidence for NO is Vals v2. Its August 13 methodology change replaced corporate-finance document analysis with a private Excel-modeling benchmark, added code migration and two legal-agent evaluations, and removed saturated SWE-bench from the coding bucket (Vals methodology). Kimi moved from a near-leader at 74.70% on v1.2 to 57.813% and tenth place on v2. That was a methodology reset, not a sudden loss of capability, but it shows that Chinese convergence was not robust across task mixes.
The boards overlap, but less than their names suggest. Artificial Analysis weights agents at 34%, coding at 24%, scientific reasoning at 24% and general capability at 18% (methodology). Vals weights finance at 54.1%, coding at 37.8% and legal work at 8.1%, based on its published GDP-weighting formula. Terminal-Bench 2.1 is the only literal component shared by both. A model can therefore reach first on Artificial Analysis through science, knowledge and broad agent performance while remaining weaker on Excel, finance or US legal work.
Release mechanics help YES. Kimi K3 was evaluated by both organizations within days of its July 16 launch, and Artificial Analysis reported that six labs had crossed 50 points by July 17 (frontier-release review). The question only needs a short overlap before a US competitor responds. But evaluation lag can also prevent overlap: a Chinese model may top one board and be displaced before the second evaluation finishes.
I combined three quantitative views:
0.68 × 0.52 × 0.98 = 34.65%.Weighting these models 40%, 35% and 25% produces 35.63%. The inputs are judgments, not fitted parameters. The model spread is more informative than the second decimal place.
The widely repeated GLM-5.3 Vals scores of 71.48% and 69.53% are not Vals Index scores. They are simple averages across the benchmark cards then present on GLM-5.3's model page: three benchmarks in the August 19 archive, and eight on the current page. GLM-5.3 has no Vals Index entry, while the canonical v2 board explicitly lists Claude Opus 5 first and Kimi K3 as the best Chinese model. Treating GLM-5.3 as the current Vals leader materially overstates the YES probability.
A second trap is the claim that GLM-4.7 previously led Vals overall. Vals used loose wording in a social post, but its December 28, 2025 archived homepage shows GPT-5.2 first overall at 64.49% and GLM-4.7 first only among open-weight models, at 56.21% and ninth overall. The apparent historical Chinese victory disappears once the open-weight and overall boards are separated.
The data that would narrow the forecast most is a timestamped public history for both leaderboards, Vals scores for GLM-5.3 and later Chinese models on all seven v2 components, and observed evaluation turnaround times for major releases. Until those exist, 35.6% is the best balance between the real near-misses and the harder same-model, two-board requirement.
Hover a data point to trace its series, or click to view the forecast generated at that time.
Signed forecast receipt
Signed Aug 31, 2026, 2:54 PM with ed25519 key preseen-prod-ed25519-20260523 and externally timestamped Aug 31, 2026, 2:54 PM.
sha256:191867453b942f...a12b637afd