指令遵循(Instruction Following)

一句话定义:模型按要求约束自己输出的能力——“写超过 400 词”、“关键词至少出现 3 次”、“用 JSON 输出”、“不要提到 X”。它看起来是最简单的一维,实际上是最容易被低估计的一维。


1. 为什么它不简单

传统评测用人工打分或 LLM 裁判,两者都不可靠。IFEval(arXiv:2311.07911)的出发点就是这个:

“Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM.”

IFEval 的解法:只用可被程序自动判定的”可验证指令”(verifiable instructions),例如”写超过 400 词""关键词至少出现 3 次”。共 25 类、约 500 个提示。

设计含义:判分与模型能力解耦——用规则判定,而不是让另一个模型来当裁判。

2. 小模型在格式遵循上未必输大模型

BFCL 榜单上的直接证据(见 工具调用):3B 的 Nanbeige4-3B-Thinking 得 51.4,高于 27B 的 Gemma-3-27b-it(29.47)、70B 的 Llama-3.3-70B(31.9)。

原因是:指令遵循是专用后训练的产物,不是参数量的产物。

3. 指令遵循高度依赖”指令是否可验证”

LangChain 在工程实践中发现:

“The most common failure pattern was that the agent wrote a solution, re-read its own code, confirmed it looks ok, and stopped.”

最常见的失败模式是——智能体写完方案、自己重读一遍代码、觉得没问题,然后就停了。它没有真正去验证代码能不能跑通。

“Teaching Agents to Write Testable Code: Agents don’t know how their code needs to be testable. We add prompting say their work will be measured against programatic tests … Forcing models to conform to testing standards is a powerful strategy to avoid ‘slop buildup’ over time.”

含义:模型的”遵循”表现,很大程度上取决于指令是否被框架以确定性方式注入并被验证,而不是模型自己记不记得住。

4. 领域规则(policy)也是指令遵循的一部分

τ-bench论文 做了一个消融:移除系统提示词里的领域策略后,gpt-4o 在 τ-airline 上从 33.2 掉到 10.8——降 22.4 个百分点。

⚠️ 论文自身表格不一致:Table 2 的 airline 基线为 35.2,Table 3 为 33.2。引用时须注明。

这条对智能体开发极其重要:模型的很多”领域能力”其实活在系统提示词里,不在参数里。

5. 后训练耦合:换了工具定义就掉分

Vivek Trivedy(LangChain)的观察:

“It shows up in ways like how changing tool logic leads to worse model performance… A truly intelligent model should have little trouble switching between patch methods, but training with a harness in the loop creates this overfitting.”

它表现为——换个工具的实现方式,模型的表现就变差。……一个真正聪明的模型,切换工具的写法应该毫无困难;但「带着某个框架一起训练」会造成这种过拟合。

6. 一个测量上的警告:别检查”路径”

Anthropic 官方的建议:

“There is a common instinct to check that agents followed very specific steps like a sequence of tool calls in the right order. We’ve found this approach too rigid and results in overly brittle tests, as agents regularly find valid approaches that eval designers didn’t anticipate. So as not to unnecessarily punish creativity, it’s often better to grade what the agent produced, not the path it took.”

人们常有一种本能——去检查智能体有没有严格按指定步骤走(比如工具调用的顺序对不对)。我们发现这种做法太死板,会让测试变得极其脆弱,因为智能体经常能找到评测设计者没预料到、但同样有效的解法。

而且他们会举出模型”违规但更好”的例子:

“Opus 4.5 solved a τ2-bench problem about booking a flight by discovering a loophole in the policy. It ‘failed’ the evaluation as written, but actually came up with a better solution for the user.”

7. 为什么它值得单独当成一维

  • 它不随参数量单调提升(第 2 节)。
  • 它对长程任务的影响被放大:Lilian Weng 引 Lin et al. 2026 的结论——用好框架的关键能力就是”及时正确调用 skills/tools + 长程指令遵循”,而这项能力对中间档模型是非单调的。
  • 它是最容易通过工程改善的一维:把指令写成可验证的、把规则注入到框架里,比换更大的模型便宜得多。

相关页面