一句话摘要:DeepMind 用 400 多个模型重新做了缩放实验,指出业界”大模型训练严重不足”——模型尺寸和训练 token 数应该等比例放大(模型每翻倍,数据也翻倍),并用 70B 的 Chinchilla 干翻了 280B 的 Gopher。
原始标题
Training Compute-Optimal Large Language Models(Hoffmann et al., DeepMind,arXiv:2203.15556,2022-03-29)
来源:https://arxiv.org/abs/2203.15556 ;HTML 全文 https://arxiv.org/html/2203.15556v1
它解决什么问题
Kaplan 2020 建议”算力涨了主要加参数”。Chinchilla 的作者质疑:这个结论建立在所有模型都训练同样多的 token 的实验设计上,等于人为限制了数据的贡献。他们要让数据量也自由变化,重新找最优解。
关键要点
1. 三种方法,同一个答案:参数与数据等比例放大
| 方法 | () | () |
|---|---|---|
| Approach 1:取训练曲线最小值 | 0.50 | 0.50 |
| Approach 2:IsoFLOP 曲线 | 0.49 | 0.51 |
| Approach 3:损失参数化建模 | 0.46 | 0.54 |
| 对照组:Kaplan et al. 2020 | 0.73 | 0.27 |
“All three approaches suggest that as compute budget increases, model size and the amount of training data should be increased in approximately equal proportions.”
三种方法都指向同一个结论——算力预算增加时,模型规模和训练数据量应该按大致相同的比例一起增加。
2. “约 20 tokens/parameter” 的来历
论文本身摘要里没有逐字出现 “20 tokens per parameter”。这个被到处引用的数字有两个出处:
- Chinchilla 自己的配置:Table 1 逐字行
Chinchilla 70 Billion 1.4 Trillion→ 70B 参数 / 1.4T tokens = 20 tokens/parameter - Epoch AI 官方的复述(脚注 3):
“According to Hoffmann et al. (2022), for dense models this is achieved by training models on around 20 tokens for each model parameter.”
Table 3 的投影逐行验证了这个比例:400M 参数对应 8.0B tokens(20.0×)、1B 对应 20.2B(20.2×)、10B 对应 205.1B(20.5×)、67B 对应 1.5T(22.4×)。
3. 核心诊断:当时的模型都”喂得太少”
“We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.”
我们研究的问题是——给定算力预算,模型该做多大、该喂多少 token。结论是当前的大模型都严重训练不足,因为大家这些年只顾着扩规模,却把数据量固定住了。
4. 70B 打败 280B
“Based on our estimated compute-optimal frontier, we predict that for the compute budget used to train Gopher, an optimal model should be 4 times smaller, while being training on 4 times more tokens.”
按我们估算的计算最优前沿,训练 Gopher 的那份算力,应该用在一个小 4 倍、但多喂 4 倍 token 的模型上。
| 模型 | 参数量 | 训练 tokens | MMLU(5-shot) |
|---|---|---|---|
| Gopher | 280B | 300B | 60.0% |
| Chinchilla | 70B | 1.4T | 67.6%(摘要写 67.5%) |
“Remarkably, Chinchilla even outperforms the expert forecast for June 2023 of 63.4% accuracy.”
Chinchilla 的成绩甚至超过了专家们对「2023 年 6 月 MMLU 达到 63.4%」的预测。
“Due to being 4× smaller than Gopher, both the memory footprint and inference cost of Chinchilla are also smaller.”
因为比 Gopher 小 4 倍,Chinchilla 的显存占用和推理成本也都更小——这是小模型路线的额外红利,也是后来产业界普遍选择过训练的动机。
重要引用(英文原文)
“By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.”
我们训练了 400 多个语言模型,规模从 7000 万到 160 亿参数,训练 token 从 50 亿到 5000 亿。结论是:要做到计算最优,模型规模和训练 token 数必须等比例放大——模型每翻一倍,训练 token 也要翻一倍。
边界与争议
- “计算最优”不等于”部署最优”:Chinchilla 优化的是训练算力固定下损失最低。但产业界真正关心的是推理成本——小模型多喂数据虽然训练不划算,推理却便宜得多。这条张力直接催生了 计算最优与过训练 里的”过训练”路线。
- 论文本身也承认模型规模有限:实验覆盖 70M–16B 参数,1.4T tokens 的 Chinchilla 是唯一的大规模验证点。
- 配比仍在被继续推翻:MiniCPM(arXiv:2404.06395)用 WSD 学习率调度器得出远高于 Chinchilla 最优的数据-模型比;Llama 3 的小模型更是被明确训练得”远超计算最优”。
与本文其他页面的关系
- 缩放定律 —— 概念页
- 计算最优与过训练 —— 为什么产业界反其道而行之
- 数据墙与数据耗尽研究 —— 当”多喂数据”这条路本身撞墙时