Chinchilla
类型:语言模型(DeepMind,2022) 参数量:70B 训练数据:1.4 万亿 tokens(= 20 tokens/parameter) 发布:2022-03-29,arXiv:2203.15556
为什么它重要
Chinchilla 本身不是最强的模型,但它是缩放定律史上最大的一次认知修正:用 70B 参数打败了 280B 的 Gopher,证明当时所有大模型都”喂得太少”。
“We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.”
当前所有大模型都严重训练不足,根因是这几年大家都在扩模型规模,却把训练数据量固定住了。
关键数字
| 模型 | 参数 | 训练 tokens | MMLU(5-shot) |
|---|---|---|---|
| Gopher | 280B | 300B | 60.0% |
| GPT-3 | 175B | — | — |
| Chinchilla | 70B | 1.4T | 67.6% |
“Remarkably, Chinchilla even outperforms the expert forecast for June 2023 of 63.4% accuracy.”
Chinchilla 的成绩甚至超过了专家们对「2023 年 6 月 MMLU 达到 63.4%」的预测——也就是说它提前半年多达到了行业预期。
“Due to being 4× smaller than Gopher, both the memory footprint and inference cost of Chinchilla are also smaller.”
它确立的规则
参数与训练 token 应等比例放大:模型每翻倍,数据也翻倍。经验值约 20 tokens/parameter(来自 Chinchilla 自身配置 70B / 1.4T,以及论文 Table 3 的投影表)。
| 参数量 | Chinchilla 最优对应的 tokens |
|---|---|
| 400M | 8.0B |
| 1B | 20.2B |
| 10B | 205.1B |
| 67B | 1.5T |
| 175B | 3.7T |
| 280B | 5.9T |
⚠️ “20 tokens per parameter” 这个短语没有逐字出现在论文摘要中,它来自 Chinchilla 自己的配置与 Table 3 的数字,以及 Epoch AI 官方的复述。
后续:这条规则被产业界修改了
Chinchilla 优化的是”训练算力固定下损失最低”。但产业界更关心推理成本,于是纷纷反其道而行之——把模型做小、多喂数据,即 计算最优与过训练 中的”过训练”。
Epoch AI 的参照:“For context, Llama 3-70B was overtrained by 10x.”
相关页面
- Chinchilla计算最优论文 —— 来源摘要
- 计算最优与过训练 —— 概念页
- 缩放定律 —— 它修正了什么