一句话摘要:Wei 等人定义了”涌现能力”——小模型没有、大模型才有的能力,因此无法通过外推小模型的表现来预测;如果涌现存在,继续放大规模还可能解锁新的能力。

原始标题

Emergent Abilities of Large Language Models(Wei et al., Google,arXiv:2206.07682,TMLR 2022)

来源:https://arxiv.org/abs/2206.07682v2


它解决什么问题

缩放定律 说损失随规模平滑下降、可以外推。但实践中人们发现,某些任务上的表现不是平滑提升,而是到了某个规模突然从”完全不会”变成”会了”。这篇论文系统化了这一现象。

关键要点

1. 定义

“Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models.”

放大模型已被证明能可预测地提升性能和样本效率。但本文要讨论一种不可预测的现象——如果某个能力在小模型上不存在、在大模型上存在,我们就称它是涌现的。

2. 涌现的思想来源(引自物理学家 P.W. Anderson 1972《More Is Different》)

“Emergence is when quantitative changes in a system result in qualitative changes in behavior.”

涌现就是「量的变化导致了质的变化」。这句话引自物理学家 P.W. Anderson 1972 年的名篇《More Is Different》,是「涌现」这个概念的思想源头。

3. 与可预测的缩放定律的对照

“In many cases, the effect of scale on performance can often be methodologically predicted via scaling laws—for example, scaling curves for cross-entropy loss have been shown to empirically span more than seven orders of magnitude… On the other hand, performance for certain downstream tasks counterintuitively does not appear to continuously improve as a function of scale, and such tasks cannot be predicted ahead of time.”

多数情况下,规模对性能的影响可以用缩放定律方法论来预测(交叉熵损失的曲线跨越了七个数量级);但另一些下游任务的性能反直觉地不随规模连续提升,这类任务无法提前预测。

4. 论文的政策含义

“The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.”

这句是”继续堆规模”路线最重要的一句理论背书。

重要引用(英文原文)

Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models.

边界与争议

  • 一年后就被强力反驳:Schaeffer 等人的《Are Emergent Abilities a Mirage?》指出,涌现是指标选择造成的人造现象——把”精确匹配准确率”这类非线性/不连续指标换成连续指标,涌现就消失了。详见 涌现能力是幻觉吗论文。
  • 更早的伏笔:BIG-bench(arXiv:2206.04615)在 2022 年就观察到”表现出突破性行为的任务往往涉及多步骤或使用了脆弱(brittle)指标”,与 Schaeffer 的结论方向一致。
  • 中立立场:目前主流看法是——“能力随规模突变”这个现象在严格意义上是指标假象,但”某些能力确实在某个规模后才变得可用”这个实践观察依然成立。见 涌现能力 概念页。

与本文其他页面的关系

  • 涌现能力 —— 概念页,含正反双方
  • 缩放定律 —— 平滑改善的那一半
  • 能力维度 —— 为什么不同维度在不同规模”解锁”