思维链(Chain-of-Thought, CoT)

一句话定义:让模型在给出最终答案之前,先把推理的中间步骤写出来。这个简单的改动,把”背得多”和”推得动”这两件事彻底分开了。


1. 起源

Wei et al.(arXiv:2201.11903):

“We explore how generating a chain of thought — a series of intermediate reasoning steps — significantly improves the ability of large language models to perform complex reasoning. … prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.”

我们研究的是让模型生成一串中间推理步骤(这就是思维链),看它能不能显著提升复杂推理能力。结果发现:给一个 5400 亿参数的模型只喂 8 个思维链示例,就能在数学应用题上达到当时的最好成绩。

更省事的版本是 Kojima et al.(arXiv:2205.11916)的 zero-shot CoT——只在提示词里加一句”Let’s think step by step”:

“increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large InstructGPT model”

2. 它为什么有效(这是重点)

直觉上我们会认为”思维链是让模型更认真”。但 知识操纵与思维链论文 的受控实验给出了更硬的机制解释:

模型能把每个人出生月份答对 100%、训了 25,000 条样本,却回答不了”这个人出生在偶数月吗”——除非让它先生成”October”这个词,再判断奇偶。

关键在于:“十月”这两个字必须先被写进上下文,模型才能对它做操作。语言模型不能”在脑子里”对参数里的知识做运算,它只能对已经生成的 token 做运算。

所以 CoT 不是”提示技巧”,它是把参数里的知识搬运到可操作的上下文里的必要步骤。

3. 两个重要边界(都有实验支撑)

边界一:训练时的 CoT 不能改善推理时的非 CoT 表现

“Including sufficient CoT samples in training does not enhance non-CoT inference”

边界二:改善”知识抽取”不改善”知识操纵”

“Improving model’s knowledge extraction don’t improve its manipulation ability”

边界三:反向检索无论怎么训练都不行

“language models cannot perform this task, regardless of training methods, data, or model size … This suggests that language models cannot be used as databases.”

这三条合起来说明:CoT 是补救手段,不是万能药。有些结构性限制(如反向检索)它救不了,工程上要靠外挂检索解决。

4. 一个交叉印证:MMLU vs MMLU-Pro

MMLU-Pro 论文(arXiv:2406.01574)发现了一个反转:

“models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions.”

这是”知识榜”与”推理榜”分离的直接证据:同样的 CoT,在纯知识榜上没用,在推理榜上有用——因为两个榜根本在测不同的东西。

5. 与推理期缩放的关系

CoT 是 推理期缩放 的基础设施:先让模型能”逐步想”,才谈得上”想更久”。

  • 单次 CoT = 一条推理路径
  • 多条 CoT + 投票 = 自一致性(self-consistency)
  • 多条 CoT + 验证器打分 = 重排(reranking)
  • 强制延长 CoT = 预算强制(budget forcing)

OpenAI o1 就是把这个思路做到了极致:

“o1 averaged 74% (11.1/15) with a single sample per problem, 83% with consensus among 64 samples, and 93% when re-ranking 1000 samples with a learned scoring function.”

o1 在每道题只采样一次时平均答对 74%(15 题里 11.1 题);用 64 个样本做投票一致时升到 83%;而用学到的打分函数对 1000 个样本重排时能达到 93%。

6. 对你的实际意义

  • 需要多步推理的任务,务必让模型把步骤写出来——不是”更认真”的问题,是机制上做不到。
  • 纯知识问答(“XX 的首都是哪里”)用 CoT 收益很小,别浪费 token。
  • 想让模型输出稳定,把 CoT 写进提示词模板,而不是指望它自己想起来。

相关页面