一句话摘要:DeepMind 用 400 多个模型重新做了缩放实验,指出业界”大模型训练严重不足”——模型尺寸和训练 token 数应该等比例放大(模型每翻倍,数据也翻倍),并用 70B 的 Chinchilla 干翻了 280B 的 Gopher。

原始标题

Training Compute-Optimal Large Language Models(Hoffmann et al., DeepMind,arXiv:2203.15556,2022-03-29)

来源:https://arxiv.org/abs/2203.15556 ;HTML 全文 https://arxiv.org/html/2203.15556v1


它解决什么问题

Kaplan 2020 建议”算力涨了主要加参数”。Chinchilla 的作者质疑:这个结论建立在所有模型都训练同样多的 token 的实验设计上,等于人为限制了数据的贡献。他们要让数据量也自由变化,重新找最优解。

关键要点

1. 三种方法,同一个答案:参数与数据等比例放大

方法()()
Approach 1:取训练曲线最小值0.500.50
Approach 2:IsoFLOP 曲线0.490.51
Approach 3:损失参数化建模0.460.54
对照组:Kaplan et al. 20200.730.27

“All three approaches suggest that as compute budget increases, model size and the amount of training data should be increased in approximately equal proportions.”

三种方法都指向同一个结论——算力预算增加时,模型规模和训练数据量应该按大致相同的比例一起增加。

2. “约 20 tokens/parameter” 的来历

论文本身摘要里没有逐字出现 “20 tokens per parameter”。这个被到处引用的数字有两个出处:

  • Chinchilla 自己的配置:Table 1 逐字行 Chinchilla 70 Billion 1.4 Trillion → 70B 参数 / 1.4T tokens = 20 tokens/parameter
  • Epoch AI 官方的复述(脚注 3):

    “According to Hoffmann et al. (2022), for dense models this is achieved by training models on around 20 tokens for each model parameter.”

Table 3 的投影逐行验证了这个比例:400M 参数对应 8.0B tokens(20.0×)、1B 对应 20.2B(20.2×)、10B 对应 205.1B(20.5×)、67B 对应 1.5T(22.4×)。

3. 核心诊断:当时的模型都”喂得太少”

“We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.”

我们研究的问题是——给定算力预算,模型该做多大、该喂多少 token。结论是当前的大模型都严重训练不足,因为大家这些年只顾着扩规模,却把数据量固定住了。

4. 70B 打败 280B

“Based on our estimated compute-optimal frontier, we predict that for the compute budget used to train Gopher, an optimal model should be 4 times smaller, while being training on 4 times more tokens.”

按我们估算的计算最优前沿,训练 Gopher 的那份算力,应该用在一个小 4 倍、但多喂 4 倍 token 的模型上。

模型参数量训练 tokensMMLU(5-shot)
Gopher280B300B60.0%
Chinchilla70B1.4T67.6%(摘要写 67.5%)

“Remarkably, Chinchilla even outperforms the expert forecast for June 2023 of 63.4% accuracy.”

Chinchilla 的成绩甚至超过了专家们对「2023 年 6 月 MMLU 达到 63.4%」的预测。

“Due to being 4× smaller than Gopher, both the memory footprint and inference cost of Chinchilla are also smaller.”

因为比 Gopher 小 4 倍,Chinchilla 的显存占用和推理成本也都更小——这是小模型路线的额外红利,也是后来产业界普遍选择过训练的动机。

重要引用(英文原文)

“By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.”

我们训练了 400 多个语言模型,规模从 7000 万到 160 亿参数,训练 token 从 50 亿到 5000 亿。结论是:要做到计算最优,模型规模和训练 token 数必须等比例放大——模型每翻一倍,训练 token 也要翻一倍。

边界与争议

  • “计算最优”不等于”部署最优”:Chinchilla 优化的是训练算力固定下损失最低。但产业界真正关心的是推理成本——小模型多喂数据虽然训练不划算,推理却便宜得多。这条张力直接催生了 计算最优与过训练 里的”过训练”路线。
  • 论文本身也承认模型规模有限:实验覆盖 70M–16B 参数,1.4T tokens 的 Chinchilla 是唯一的大规模验证点。
  • 配比仍在被继续推翻:MiniCPM(arXiv:2404.06395)用 WSD 学习率调度器得出远高于 Chinchilla 最优的数据-模型比;Llama 3 的小模型更是被明确训练得”远超计算最优”。

与本文其他页面的关系

  • 缩放定律 —— 概念页
  • 计算最优与过训练 —— 为什么产业界反其道而行之
  • 数据墙与数据耗尽研究 —— 当”多喂数据”这条路本身撞墙时