一句话摘要:主流的 n-gram 去污方法可以被”改写、翻译”简单绕过(检测 F1 = 0);只要把测试集的改写版本喂进训练,一个 13B 模型就能在 MMLU 上刷到 88.5–89.9%、GSM8K 刷到 86.7–95.3%——与 GPT-4 同档。
原始标题
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples(Yang, Chiang, Zheng, Gonzalez, Stoica,arXiv:2311.04850,2023-11-07)
来源:https://arxiv.org/abs/2311.04850
⚠️ 编号勘误:常见误引为 2311.09783。2311.09783 是另一篇《Investigating Data Contamination in Modern Benchmarks for Large Language Models》(Deng et al., NAACL 2024),两篇都存在且都可用,但本页讲的是 2311.04850。
它解决什么问题
各家厂商都说”我们做了去污(decontamination)“。这篇论文要检验:这些去污手段到底管不管用?
关键要点
1. 改写法可以完全绕过 n-gram 去污
“While most data decontamination efforts apply string matching (e.g., n-gram overlap) to remove benchmark data, we show that these methods are insufficient, and simple variations of test data (e.g., paraphrasing, translation) can easily bypass these decontamination measures.”
主流去污用的是字符串匹配(n-gram 重叠),我们证明这些方法不够用——对测试题做点改写或翻译,就能轻松绕过。
| 检测方法 | 在改写样本上的 F1 |
|---|---|
| n-gram 重叠 | 0(完全失效) |
| LLM decontaminator | 0.94–1.00 |
2. 13B 模型刷分幅度(论文 Figure 3)
| 榜单 | 微调前基线 | 用改写样本微调后 |
|---|---|---|
| MMLU | 约 45–54% | 88.5–89.9% |
| GSM-8K | 约 15–29% | 86.7–95.3% |
| HumanEval(CodeLlama pass@1) | 约 33–36% | 67.7–81.1% |
“Furthermore, we demonstrate that if such variation of test data is not eliminated, a 13B model can easily overfit a test benchmark and achieve drastically high performance, on par with GPT-4.”
如果不清除这些变体,一个 130 亿参数的小模型就能轻易在测试榜上刷到和 GPT-4 齐平的水平。
3. 真实预训练语料里已经存在的污染比例
“in pre-training sets such as RedPajama-Data-1T and StarCoder-Data, we identified that 8-18% of the HumanEval benchmark overlaps. Interestingly, we also find such contamination in synthetic dataset generated by GPT-3.5/4, suggesting a potential risk of unintentional contamination.”
在 RedPajama-Data-1T 和 StarCoder-Data 这类预训练集里,HumanEval 有 8–18% 的重叠;连 GPT-3.5/4 生成的合成数据里也发现了污染,提示存在无意识污染的风险。
| 数据集 | HumanEval 重叠率 |
|---|---|
| The Stack | 18.9% |
| StarCoder-Data | 15.9% |
| RedPajama-Data-1T | 8.53% |
| CodeAlpaca(GPT-3.5 合成) | 12.8% |
| MATHInstruct | 15.4% |
4. 作者的呼吁
“We urge the community to adopt stronger decontamination approaches when using public benchmarks. Moreover, we call for the community to actively develop fresh one-time exams to evaluate models accurately.”
我们呼吁社区在公开基准上采用更强的去污方法,并主动开发「全新的一次性考试」来准确评估模型——这也是后来 GSM1k、 Humanity’s Last Exam 这类榜单的思路来源。
重要引用(英文原文)
“We validate such observations in widely used benchmarks such as MMLU, GSK8k, and HumanEval.”
我们在 MMLU、GSK8k、HumanEval 这些被广泛使用的基准上验证了上述观察。(GSK8k 是原文写法,即 GSM-8K。)
补充:现有治理手段基本都无效(arXiv 2503.16402,ICML 2025)
“Extensive experiments with 10 LLMs, 5 benchmarks, 20 BDC mitigation strategies, and 2 contamination scenarios reveal that no existing strategy effectively balances fidelity and contamination resistance. No semantic-preserving strategy yields a significant improvement in resistance over the vanilla case … while semantic-altering strategies sacrifice fidelity for resistance.”
我们在 10 个大语言模型、5 个基准、20 种 BDC 缓解策略和 2 种污染场景下做了大量实验。结果显示:现有没有任何一种策略能同时兼顾「保真度」和「抗污染性」——不改变语义的策略在抗污染上几乎没有提升,而改变语义的策略则是牺牲保真度来换抗污染性。
边界与争议
- 这是”刻意污染”的实验(故意拿改写过的测试集去微调),不等于真实训练流程。它证明的是上限风险,不是”所有模型都在作弊”。
- 对照数据:GSM1k(arXiv 2405.00332)用 1,250 道全新人类出题的镜像题测真实污染,发现 Gemini/GPT/Claude 等前沿模型几乎无掉分,而 Phi、Mistral 系列掉约 10%、最多 13%。说明污染程度因厂商而异。
- 结论的正确用法:不是”跑分都不能信”,而是”公开的静态榜单天然会随时间失效,越流行的榜越容易被污染”。
与本文其他页面的关系
- 评测污染 —— 概念页
- 排行榜幻象论文 —— 污染之外,还有主动刷榜
- 按题目发布时间切分(LiveCodeBench 的做法)是主流的工程解法,见 BFCL 与 Terminal-Bench