工具调用(Tool Use / Function Calling)

一句话定义:模型输出一个结构化的”我要调用这个函数、参数是这些”的请求,由外部系统真正执行,再把结果喂回模型。这是模型从”说话”走向”做事”的接口。


1. 它包含四个可学习的决策

Toolformer(arXiv:2302.04761)最早把工具调用形式化:

“a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction.”

Toolformer 把「用工具」拆成了四个可学习的决策——调哪个 API、什么时候调、传什么参数、怎么把返回值融进后面的预测。这四个决策基本就是工具调用能力的全部内容。

决策失败时会怎样
选哪个工具用了不相关的工具
什么时候调该调不调、不该调乱调
传什么参数参数类型/取值错误
如何把结果并入后续生成拿到结果却不会用

2. 它是训练出来的,不是参数涌出来的(最重要的一条)

ToolLLM论文(arXiv:2307.16789)的原话:

“they remain significantly limited in tool-use capabilities … The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain.”

开源模型在工具使用能力上明显受限,也就是「用外部工具(API)去完成人类指令」这件事。原因在于:当时的指令微调主要聚焦在基础语言任务上,把工具使用这个领域整个忽略了。

2026 年的实证支持这个判断——在 BFCL V4 榜单上(Last Updated 2026-04-12):

模型Overall Acc
Nanbeige4-3B-Thinking(3B)51.4
xLAM-2-3b-fc-r(3B,专用工具模型)41.22
Phi-4(14B)28.79
Gemma-3-27b-it(27B)29.47
Llama-3.3-70B(70B)31.9
Llama-4-Scout-17B-16E28.13

3B 的模型把 70B 的模型甩在后面——参数量与工具调用能力不单调。这是”参数大 ≠ 工具调用强”最直接的证据。

3. 现在怎么测:BFCL 的重心已经转移

BFCL V4(2025-07 发布)的总分构成(官方逐字):

Overall Score = Agentic(40%) + Multi-Turn(30%) + Live(10%) + Non-Live(10%) + Hallucination(10%)

官方博客原话的含义:单轮函数调用已经饱和、无法区分前沿模型,所以榜单把 70% 的权重挪到了跨轮次保持状态、搜索、记忆、以及知道什么时候不该调用。

“The single hardest capability is abstention: the Hallucination category rewards a model for correctly refusing to call a function when no available tool fits, and tool-tuned models are systematically biased toward calling something.”

最后半句非常重要:专门调过工具调用的模型,会倾向于”总得调点什么”——这是个系统性的偏置。

4. 稳定性:单次跑分会骗人

τ-bench论文(arXiv:2406.12045)提出的 pass^k 揭穿了这个问题:

gpt-4o 在 τ-retail 上 pass^1 = 61.2%,但 pass^8 < 25%。

详见 一致性指标pass的k次方:“能做成一次”和”每次都能做成”是两种能力。

5. 工具调用能力的一半在模型外

这是本概念最容易被忽略的部分。Anthropic 官方的说法:

“When we evaluate ‘an agent,’ we’re evaluating the harness and the model working together.”

Terminal-Bench 2.0 官方榜给出的数字:

  • Claude Opus 4.6:Meta-Harness 76.4% vs Claude Code 58.0%(差 18.4 点)
  • GPT-5.3-Codex:LemonHarness 84.5% vs Terminus 2 64.7%(差 19.8 点)

详见 模型与框架分责 与 Terminal-Bench与Harness分责。

而 Lilian Weng 引 Lin et al. 2026 的研究,把这件事拆成了两个独立的轴:

“harness-updating refers to the capability of producing useful harness edits and harness-benefit denotes the capability of utilizing the updated harness… a range of model of different sizes and core intelligence, from Qwen3.5-9B to Claude Opus 4.6, were observed … to show similar harness updating capability; the 9B harness proposer is able to write a skill procedurally isomorphic to Opus. To best utilize a harness, a model needs to invoke skills/tools correctly and timely and be good at long-horizon instruction following.”

结论:写框架的能力跨模型尺寸基本持平;真正分化的是”用好框架的能力”——即”及时正确地调用工具 + 长程指令遵循”。

6. 下一步:智能体强化学习(agentic RL)

2025–2026 年的方向不再是模仿数据,而是在真实可执行环境里做强化学习:

  • ToolVerse(arXiv:2607.15660)从近 400 个真实 MCP(MCP 模型上下文协议)中构建约 4,500 个工具的可执行训练环境
  • ATLAS(arXiv:2603.06713)针对小模型在大工具空间里的三个失效模式:“eager tool loading saturates context, execution errors compound over time, and sparse rewards limit learning”

相关页面