推理期缩放(Test-time Scaling / Inference-time Scaling)
一句话定义:不改模型权重,只在”回答问题的时候多花算力”——让模型想更多步、生成多个候选再挑、或者边做边验证。这是继”把模型训大”之后的第二条提升曲线。
1. 两个旋钮
Raschka(2026-01-24)的表述最简洁:
“I think this figure … nicely captures the idea behind the two knobs we can use to improve LLMs. We can spend more resources during training (more data, bigger models, more or longer training stages) or inference.”
这张图很好地概括了改进大模型可用的两个旋钮:你可以把资源花在训练阶段(更多数据、更大模型、更长训练),也可以花在推理阶段。
“Actually, in practice, it’s even better to do both at the same time: train a stronger model and use additional inference scaling to make it even better.”
OpenAI 在 o1 发布时(2024-09-12)说的是同一件事:
“We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). The constraints on scaling this approach differ substantially from those of LLM pretraining.”
OpenAI 发现 o1 的表现同时随两件事提升——训练时多花强化学习算力,以及推理时多花时间思考。而且这两条路线的瓶颈和预训练完全不同:预训练受数据量限制,这两条不受。
2. 量化证据:想更久能换多少
| 证据 | 数字 | 来源 |
|---|---|---|
| o1 在 AIME 2024 上 | 单次采样 74.4% → 64 样本投票 83.3% → 1000 样本重排 93%(13.9/15) | OpenAI o1 官方博客 |
| 计算最优的测试时分配 vs 朴素 best-of-N | 效率提升 >4× | Snell et al., arXiv:2408.03314 |
| FLOPs 匹配下,小模型用测试时算力 | 可以超过 14× 大的模型 | 同上 |
| s1-32B 用”预算强制”延长思考 | AIME24 从 50% → 57% | arXiv:2501.19393 |
| Raschka 书中对基座模型的提升 | 约 15% → 52% | Raschka 博客 2026-01-24 |
⚠️ 14× 那条有严格的限定条件:“on problems where a smaller base model attains somewhat non-trivial success rates”——小模型本来就有一定成功率的题目才成立,不是万能的。
3. 常见手段
-
多数投票 / 自我一致性(self-consistency):生成 k 个解,选出现最多的答案。
-
验证器重排(verifier reranking):用过程奖励模型(PRM)给每个候选打分,选最高的。
-
预算强制(budget forcing):s1 论文的做法——模型想结束思考时,强行追加 “Wait” 让它再检查一遍:
“forcefully terminating the model’s thinking process or lengthening it by appending ‘Wait’ multiple times to the model’s generation when it tries to end. This can lead the model to double-check its answer, often fixing incorrect reasoning steps.”
-
智能体式工具循环:把”想”换成”做”——调用工具、看结果、再想。见 工具调用。
4. s1 的极端案例:1,000 条数据能走多远
“First, we curate a small dataset s1K of 1,000 questions paired with reasoning traces relying on three criteria we validate through ablations: difficulty, diversity, and quality. Second, we develop budget forcing…”
结果:只用 1,000 条精选问答对 Qwen2.5-32B-Instruct 做微调 + 预算强制,在竞赛数学题上最高超出 o1-preview 27%。
这条对”参数决定一切”是个有力的反例:同一份权重,光改推理时的策略,就能拉开 27 个百分点的差距。
5. 它和”参数量”的关系
- 互补,不是替代:Snell et al. 的论文标题已经把话说完——“Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters”(在给定条件下更有效,不是永远更有效)。
- 难度依赖:测试时算力的效果强烈依赖题目难度,所以最优策略是”按题分配算力”——简单题少想,难题多想。
- 它改变了成本结构:参数量决定的是固定成本(训练 + 部署),推理期缩放决定的是可变成本(每次回答花多少 token)。这也意味着同一个模型的”能力”不再是固定值,而取决于你愿意花多少钱。
相关页面
- 缩放定律 —— 第一条曲线
- 数据墙 —— 为什么需要第二条曲线
- 能力维度 —— 推理期缩放主要提升哪一维