What 106,134 arena votes can and cannot order
A leaderboard is a ranked list, and the ranked list is what everyone reads. On the public LMArena sample, the votes separate one neighbouring pair out of 54.
Arena leaderboards turn pairwise votes into a ranked list, and the ranked list is what everyone reads. But a rank is a claim, and the votes behind it support some of those claims and not others. Two models a hair apart may be tied. A board can look decisive because a handful of people voted a great many times. So I took the public LMArena preference sample (106,134 human votes over 55 models), fitted the same Bradley–Terry model the boards use, and asked which parts of the published order the evidence actually carries.
One separated pair out of 54
Rank 1 versus rank 2, rank 2 versus rank 3, and so on: 54 neighbouring comparisons. Judged one at a time, five of them are separated. Controlled for together, the right way to read 54 simultaneous claims, only one is. Only the top model is distinguishable from its neighbour. The other 54 models form a single block that these votes cannot order.
That is not the same as saying the published order is wrong. These votes do not establish it, a more boring claim and the one the data supports.
Closing the median unresolved gap would take roughly 8.9 million votes at this design, against the 106 thousand collected. For most neighbouring pairs, the right reading is that they are tied and will stay tied. The median model’s 95% rank interval spans 7 ranks; gemma-2-9b-it-simpo, with 468 battles, spans 16.
Position matters, and it does not matter
This is my favourite result in the audit, because it refuses to resolve into a headline.
The model shown first wins 49.53% of 65,418 decisive battles. The interval on the fitted first-slot advantage, [−0.021, −0.001] in log-odds, excludes zero. At this sample size the effect is real, and anyone reporting “no position bias” would be wrong. The entire interval also sits inside the ±0.05 band I treat as immaterial, and fitting the advantage out reverses no separated pair, though four models do shift rank inside the tied block.
Detectable and worth acting on are different questions. With 106,134 votes you can detect effects far too small to change any decision, and a leaderboard needs an answer to the second question. Reporting only the p-value answers the first and implies the second.
The checks that came back clean
An audit that only ever finds problems is not an audit, so the negative results matter as much:
- Voter clustering widens standard errors by 1.03×. The busiest 1% of the 49,382 voters cast 16% of all votes, so the concern is well founded. On this data, though, the usual battle-level bootstrap is close enough. That was checked, and it could have come out the other way.
- 38.4% of votes are ties, and the ranking survives both dropping them and modelling them properly (Rao–Kupper theta 2.35) instead of splitting them down the middle.
- Preference cycles land about where the model expects. 3.5% of the 14,164 decided triads are cyclic, against 4.3% predicted. One strength number per model is enough to describe who beats whom here; there is no rock-paper-scissors structure being flattened.
The same six models, judged by people and by GPT-4
MT-Bench publishes both halves of the same comparison: 3,355 votes from 65 people, and 2,400 verdicts from GPT-4 acting as a judge, on the same models and the same questions. Auditing both is the cleanest available test of what judge bias does to a ranking.
GPT-4 changes its mind on 15.8% of pairs when the two answers swap places. That figure is not estimated: the split records those cases explicitly. It matches the human majority on 71.3% of the 1,612 comparisons where the humans reached one. Both numbers look disqualifying.
Then it produces the human leaderboard, rank for rank. Kendall tau between the two orders is +1.00.
Judge noise that looks fatal one comparison at a time can average out of a six-model ranking. That cuts in both directions, and it is worth being careful about which direction you invoke it in: it is a reason not to condemn a leaderboard because its judge is noisy, and equally a reason not to trust a leaderboard because its judge agrees with people on average. The aggregate and the individual verdict are answering different questions.
The judge’s verdicts are also correlated within question: resampling at the question level widens its intervals by 1.48×. The human votes, spread over 65 annotators, show no annotator carrying the board.
How I read a leaderboard now
Mostly by asking what it would take to move. If the gap between two models is smaller than the interval, they are tied, and the ordering between them is a rendering artefact of having to print a list in some order. That is not a criticism of arenas, which are doing something hard and doing it in public. It is a criticism of reading a ranked list as though ranks were the measurement.