一句话摘要:论文明确点出开源模型工具调用弱的根因——“当前的指令微调主要聚焦基础语言任务,忽略了工具使用领域”。这句话是”参数大 ≠ 工具调用强”最直接的一手论据。
原始标题
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs(Yujia Qin 等,清华大学,arXiv:2307.16789,v1 2023-07-31)
来源:https://arxiv.org/abs/2307.16789
它解决什么问题
2023 年,ChatGPT 能熟练调用外部工具,而同等甚至更大参数量的开源模型做不到。作者要搞清楚:差距到底来自模型规模,还是来自训练数据的覆盖面?
关键要点
1. 根因诊断(本页最核心的一句)
“Despite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to fulfill human instructions. The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain. This is in contrast to the excellent tool-use capabilities of state-of-the-art (SOTA) closed-source LLMs, e.g., ChatGPT.”
这句话的含义:工具调用是后训练(post-training)阶段专门喂出来的能力,不会随预训练参数量自动出现。
2. 数据怎么造的(三阶段)
“(i) API collection: we collect 16,464 real-world RESTful APIs spanning 49 categories from RapidAPI Hub; (ii) instruction generation: we prompt ChatGPT to generate diverse instructions involving these APIs, covering both single-tool and multi-tool scenarios; (iii) solution path annotation: we use ChatGPT to search for a valid solution path (chain of API calls) for each instruction.”
数据构造分三步——(i) 从 RapidAPI Hub 收集了 49 个类别、共 16464 个真实 API;(ii) 让 ChatGPT 围绕这些 API 生成各种指令,覆盖单工具和多工具场景;(iii) 再让 ChatGPT 为每条指令搜出一条可用的调用链。
3. 推理侧的改进:DFSDT
“To enhance the reasoning capabilities of LLMs, we develop a novel depth-first search-based decision tree (DFSDT) algorithm. It enables LLMs to evaluate multiple reasoning traces and expand the search space.”
为了增强推理能力,他们提出了基于深度优先搜索的决策树(DFSDT),让模型能同时评估多条推理路径、扩大搜索空间,而不是一条道走到黑。
4. 结果
“Based on ToolBench, we fine-tune LLaMA to obtain an LLM ToolLLaMA … ToolLLaMA demonstrates a remarkable ability to execute complex instructions and generalize to unseen APIs, and exhibits comparable performance to ChatGPT.”
关键:ToolLLaMA 是 7B 量级的模型,靠专门的工具调用数据追平了 ChatGPT 的工具使用表现。
重要引用(英文原文)
“The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain.”
原因在于——当前的指令微调主要聚焦在基础语言任务上,把工具使用这个领域整个忽略了。
前史:Toolformer(arXiv 2302.04761,Meta AI)
更早的 Toolformer 把工具调用拆成四个可学习决策:
“We introduce Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API.”
过滤准则(自监督):只保留能降低后续 token 预测损失的调用,即 。基线模型是 GPT-J(6B)。
边界与争议
- 论文是 2023 年的,工具调用能力此后已被大规模内化进前沿模型的后训练流程;但”这是训练产物而非规模产物”的机制判断没有过时。
- 2026 年的新证据支持同一结论:在 BFCL V4 榜单上,3B 的 Nanbeige4-3B-Thinking 得 51.4,高于 27B 的 Gemma-3-27b-it(29.47)、70B 的 Llama-3.3-70B(31.9)、14B 的 Phi-4(28.79)——参数量与工具调用能力不单调。
- 下一步是智能体强化学习(agentic RL):2025–2026 年的工作(如 ToolVerse、ATLAS)直接在可执行工具环境里做强化学习,而不是靠模仿数据。