一句话摘要:Epoch AI / Villalobos 等人估算,人类公开文本数据的有效存量约 300 万亿 token,按现有趋势将在 2026–2032 年间被用尽;而产业界为了降低推理成本纷纷”过训练”,这会把耗尽时间提前到 2025–2027 年。
原始标题
Will we run out of data? Limits of LLM scaling based on human-generated data(Villalobos et al.,arXiv:2211.04325,v1 2022-10-26,v2 2024-06-04)
官方更新版(Epoch AI 出版物页,2024-06-06):https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data
它解决什么问题
Chinchilla计算最优论文 之后的产业共识是”多喂数据”。但互联网上的高质量人类文本是有限的存量,不是可以无限开采的流量。这篇论文要算清楚:这个存量有多大,什么时候会被用完。
关键要点
1. 存量估算
“We find that the total effective stock of human-generated public text data is on the order of 300 trillion tokens, with a 90% confidence interval of 100T to 1000T. This estimate includes only data that is sufficiently high-quality to be used for training, and accounts for the possibility of training models for multiple epochs.”
人类生成的公开文本,剔除低质量内容、并考虑多轮重复训练之后,有效存量约为 300 万亿 token(90% 置信区间 100T–1000T)。
2. 耗尽时间
“Our 80% confidence interval is that the data stock will be fully utilized at some point between 2026 and 2032.”
我们有 80% 的把握认为,这批数据会在 2026 到 2032 年之间的某个时点被完全用尽。
“If models are trained compute-optimally, there is enough data to train a model with 5e28 floating-point operations (FLOP), a level we expect to be reached in 2028.”
如果按计算最优的方式训练,现有数据足够训出一个 5e28 FLOP 规模的模型——而我们预计这个规模要到 2028 年才会达到。也就是说,数据还没立刻卡死这条路。
3. 过训练(over-training)的经济学动机——最关键的一段
“We develop a simplified model of revenues and costs, and solve it to find the scaling policy that would maximize AI developers profits. Depending on how demand for AI inference changes with the performance of the model, we find it might make sense to overtrain models up to 100x. … If models are overtrained by a modest factor of 5x, the stock of data will be fully used by 2027, but if they are overtrained by 100x, the stock of data will be fully used by 2025. For context, Llama 3-70B was overtrained by 10x.”
为什么厂商愿意过训练?因为训练是一次性成本,推理是持续成本。把模型做小、多喂数据,虽然训练时不划算,但每次服务用户都更便宜。
4. 术语定义(脚注原文)
过训练(overtrained)的定义:
“Overtrained models are those that are trained on more data than what is prescribed by compute-optimal scaling laws. Holding training compute constant, overtrained models are more efficient during inference (because they have fewer parameters), at the cost of more data usage and somewhat lower performance.”
过训练的模型 = 喂了超过「计算最优」所需的数据量。在训练算力相同的前提下,它们因为参数更少而推理更便宜,代价是消耗更多数据、性能略低。
5. 固定数据下的收益上限
“Assuming the dataset size is fixed at 300T tokens, the gains that can be achieved by undertraining models eventually plateau at a level that is equivalent to ~2 additional orders of magnitude of compute-optimal scaling.”
如果把数据集固定死在 300T token,那么「少训练一点」能带来的收益最终会触顶——天花板大约相当于再多两个数量级的计算最优缩放。
6. 作者给出的三条出路
“While we cannot predict which innovations will succeed, we identify three categories that seem especially relevant: synthetic data, learning from other modalities of data, and data efficiency improvements.”
我们无法预测哪些创新会成功,但认为三类方向特别相关:合成数据、从其他模态学习、以及提升数据利用效率。
重要引用(英文原文)
“Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.”
我们的发现表明:如果当前的大模型发展趋势持续下去,那么在 2026 到 2032 年之间,模型训练所用的数据集规模就会追平公开人类文本数据的总存量;如果模型被「过训练」,这个时点还会再早一些。
边界与争议
- 这是 2022 年的估算,2024 年更新过一版。2025 年 Epoch AI 自己的《Can AI scaling continue through 2030?》口径明显更乐观:把多模态与合成数据算进来后,估计 2030 年可用数据当量达 400 万亿–2 亿亿 tokens。
- “过训练倍数”是本页最容易被误用的数字:只有 Llama 3-70B 的 10× 是 Epoch AI 原文给出的;其他模型的倍数都是后人按 (tokens/param) ÷ 20 推算的,不是原文结论。
- 最新的实时数据见 Epoch AI 实体页:截至 2026-08,前沿模型训练算力仍在以每年约 5 倍的速度增长,没有明显放缓迹象。
与本文其他页面的关系
- 数据墙 —— 概念页
- 计算最优与过训练 —— 为什么”多喂数据”成了产业标配
- 推理期缩放 —— 数据不够时,算力改花在”想更久”上