How stable are AI leaderboards? One refresh moved 135 model scores
By Data KhosrowshahiParody pen name · View profile135 model scores changed in one August 29 refresh; 129 of them drifted upward together by at most 0.22 points

Most single-day movement on AI leaderboards is the scorer, not the models. In the August 29, 2026 refresh of our model score tracking, 135 models had their composite scores (one number aggregating many benchmark results) updated at once, and 129 of them drifted upward together by at most 0.22 points. The drifting models came from 29 different makers: Google, xAI, Alibaba, Anthropic, Moonshot, and two dozen more, all nudged the same direction on the same morning. Twenty-nine competitors do not ship improvements on the same morning; a change to the scoring method moves them all at once. One real story hid inside the same refresh: a single model, Qwen3.8-Flash-Next, fell 6.22 points while nothing else moved more than 0.22 in either direction. Reading a leaderboard means telling those two movements apart: the scoring method shifting under everyone, versus one model being re-measured. The difference is invisible if you only quote one model’s number.
Key Takeaways
- One routine refresh changed 135 of 24,840 tracked model score rows in a day.
- 129 of the 135 changes were small upward drifts, at most 0.22 points, spread across 29 different makers: the signature of recalibration, not of synchronized capability gains.
- The one large move confirms the test: a single model fell far outside the band on its own, which reads as a re-measurement rather than a change in the scoring method.
What changed in the August 29 refresh?
The refresh profile has two parts. A tight band of small gains, median just over a tenth of a point, covering nearly every model that changed. And one outlier: Qwen3.8-Flash-Next fell on its own, by the margin the figure draws, while the next-largest decline in the entire refresh was 0.12. The band and the outlier are different kinds of events.
The band moved models regardless of maker, size, or age, which no plausible set of independent model updates produces. The outlier moved one model by fifty times the median, which no rounding pass produces. Reading the band as recalibration and the outlier as a re-measurement of one model is our inference from the shape of the score changes, not a disclosure from any scorer.
How do you tell recalibration from capability change?
Three checks work on any leaderboard, ours included:
- Cross-maker sync. Models from 29 unrelated companies moving the same direction on the same day points at the scorer, not the models.
- Breadth. A pass touching 135 scores at once is systematic. A capability release moves one model and its close variants.
- Magnitude sorting. Recalibrations produce many small, same-signed changes. Re-measurements of one model produce a few large, isolated ones.
The same lens applies to third-party composites. Surge’s Tuesday Work Index runs for DeepSeek V4 Pro two days earlier and for Qwen 3.8 Max the week before are single-model measurements against published criteria, which is the interpretable kind of move. The uninterpretable kind is a composite score that changes without saying whether its definition changed. Our ground-truth research raises the same trust problem about verifiers and judges.
“V4 Pro scores 59.7 on the Tuesday Work Index”
Is the AI benchmark catalog itself stable?
No. The same day’s snapshot added 11 benchmarks, taking the catalog from 3,079 to 3,090, and the tracked environment count moved from 100 to 101, its first change in our daily observation window. Since the June census counted 3,029 benchmarks against 99 environments, the stock has grown by 61 benchmarks and 2 environments. The benchmark catalog grows by the day; the environment count moves by ones.
What this means
Treat any single-day leaderboard move as a measurement event until the surrounding score changes say otherwise. Cross-maker synchronization is the fastest recalibration test. A composite score is only as stable as its definition, and the August 29 changes read as a definition change, not a capability change.
FAQ
Why would leaderboard scores get recalibrated at all?
Composite scores aggregate many benchmark results through weights and normalization, and those inputs change constantly: benchmarks get added or corrected, leaderboards update, aggregation rules improve. Each such change reprices every model it touches, which is how 135 scores can move in one day while almost none of the models did anything.
Does recalibration make a leaderboard untrustworthy?
Recalibration does not make a leaderboard untrustworthy; it makes it an instrument, and instruments need recalibration. The trust question is visibility: a leaderboard that shows what moved and by how much supports interpretation, while one that shows only today’s number invites reading every move as capability.
How many models, benchmarks, and RL environments does rlresearch.ai track?
rlresearch.ai tracked 24,840 model score rows, 3,090 benchmarks, and 101 RL environments as of August 29, 2026. Daily observation of day-over-day changes began in mid-August 2026.