一句话摘要:Meta 在 Llama-4 发布前私下测试了 27 个变体、只公开最好的那个;Google 和 OpenAI 各自拿走了对战榜 19.2% 和 20.4% 的数据,而 83 个开源权重模型加起来只占 29.7%——对战榜的分数不只反映模型质量,还反映厂商的数据特权。

原始标题

The Leaderboard Illusion(arXiv:2504.20879,NeurIPS 2025 Datasets & Benchmarks)

来源:https://arxiv.org/html/2504.20879


它解决什么问题

对战榜(如 LMArena)被认为是”最难刷”的评测,因为题目是实时的人类提问、没有静态测试集可背。这篇论文要检验:它真的刷不动吗?

关键要点

1. 私有多版本测试(best-of-N)

“We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release.”

数学表达:如果 是第 k 个变体的估计实力,那么 best-of-N 策略导致 —— 光靠”多试几次只报最好的”,分数就会被系统性抬高。

2. 数据访问不对称

“Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data.”

Google 和 OpenAI 分别拿走了平台上约 19.2% 和 20.4% 的数据;相比之下,83 个开源权重模型加起来只拿到约 29.7%。数据获取本身就是不平等的。

3. 用对战榜数据训练能大幅提分

“With conservative estimates, we show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution.”

按保守估计我们证明,能拿到 Chatbot Arena 的数据会带来巨大好处——哪怕只多一点点数据,也能在 ArenaHard 上带来最高 112% 的相对性能提升。而 ArenaHard 这个测试集,本身就来自 Arena 的数据分布。

关键补充:“This improvement does not translate to out-of-distribution performance on benchmarks such as MMLU.”

4. 模型下架不透明

“Most models (205 out of 243) are silently deprecated… Open-weight and open-source models are more likely to be deprecated.”

大多数模型(243 个里有 205 个)是被静默下架的——不打招呼、不发公告,直接就从榜上消失了。而开源权重的模型被下架的概率更高。

摘要中的量化:静默下架的模型中 64% 是 open-weight 或 open-source。

5. 卷首题词(适合做文章引子)

“Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes. — Charles A. E. Goodhart”

任何被观察到的统计规律,一旦有人出于控制目的对它施压,就会趋于崩塌。这是古德哈特定律,论文用它作卷首题词——榜单一旦成为目标,就不再是好指标。

重要引用(英文原文)

“We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. … Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality.”

我们发现,这种不公开的私测做法让少数厂商受益——他们能在公开发布前测试多个变体,并且想撤分就撤分。……这些动态加在一起,结果是模型在过拟合 Arena 特有的规则,而不是在提升通用质量。

边界与争议

  • LMArena 官方强烈反驳:联创 Ion Stoica 称该论文 “full of inaccuracies” 和 “questionable”,并回应”Meta 对我们政策的理解不符合我们对模型提供方的期望”。
  • 本页部分正文数字(2M battles、42 providers、243 models;arena 数据占比 0%→70% 使 ArenaHard 胜率 23.5%→49.9%;12 月提示在次年 1 月逐字重复 7.3%)来自二次文献综述,未逐字核对 arXiv 正文。摘要中的”27 个私有变体”已核实。

与本文其他页面的关系

  • 偏好对战榜 —— 概念页,讲 Elo/Bradley-Terry 机制
  • Arena风格控制实验 —— 榜单官方自己做的偏见研究
  • 评测污染 —— 静态榜的污染问题