模型与框架分责(Model-Harness Responsibility Split)

一句话定义:你看到的任何一个”智能体跑分”,测的都是”模型 + 框架”这个组合,不是模型单独的能力。分清哪部分归模型、哪部分归框架,是读智能体评测的第一道门槛。


1. 公式

“Agent = Model + Harness. If you’re not the model, you’re the harness.”(Addy Osmani / Vivek Trivedy)

智能体 = 模型 + 框架。只要你做的不是模型本身,那你做的就是框架。

Lilian Weng 的定义更完整:

“A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.”

框架就是围在基座模型外面的那一层系统。它负责安排执行流程,并决定模型怎么思考和规划、怎么调用工具和行动、怎么感知和管理上下文、怎么保存产物、怎么评估结果。

框架提供的四样模型天生没有的东西:

“Out of the box they cannot: Maintain durable state across interactions / Execute code / Access realtime knowledge / Setup environments and install packages to complete work. These are all harness level features.”

模型开箱即用地做不到这几件事:跨多轮对话保持持久状态、执行代码、获取实时信息、搭建环境并安装软件包。这些全都是框架层提供的功能。

2. 分责有多重要:18–20 个百分点的差距

Terminal-Bench 2.0 官方榜(抓取 2026-08-31):

模型harness准确率
Claude Opus 4.6Meta-Harness76.4% ± 2.4
Claude Opus 4.6Claude Code(自家)58.0% ± 2.9
GPT-5.3-CodexLemonHarness84.5% ± 2.6
GPT-5.3-CodexTerminus 264.7% ± 2.7

LangChain 的对照实验更干净(模型固定,只改框架):

“We used a simple recipe to iteratively improve deepagents-cli … 13.7 points from 52.8 to 66.5 on Terminal Bench 2.0. We only tweaked the harness and kept the model fixed, gpt-5.2-codex.”

我们只改框架、模型一个字没动(固定为 gpt-5.2-codex),分数就从 52.8 涨到 66.5,涨了 13.7 个点。

“Our coding agent went from Top 30 to Top 5… We only changed the harness.”

我们的编程智能体从第 30 名开外升到了前 5 名,而唯一的改动就是换了框架。

3. 一句常被引用但必须加限定的话

“A decent model with a great harness beats a great model with a bad harness.”(Addy Osmani)

限定一:框架不是越强越好,要跟模型匹配。 CORE-Bench 的对照:

“when we ran Claude Opus 4.5 using Claude Code, it scored 78%, nearly double the 42% we reported using our standard CORE-Agent scaffold.”

换上 Claude Code 之后,Opus 4.5 的分数从 42% 涨到 78%,几乎翻倍。这说明框架和模型是配套的,不存在通用最优框架。

“This gap was much smaller for other models: For Opus 4.1, CORE-Agent outperforms Claude Code by almost 10 percentage points.”

作者的假设:“the Claude 4.5 series of models is much better tuned to work with Claude Code”;或”the lower-level instructions in CORE-Agent, which worked well for less capable models, stop being effective (and hinder the model’s performance) for more capable models”

限定二:模型必须”够格”,框架才能起正面作用。 Lilian Weng 引 Zelikman 2023:

“STOP improved mean downstream performance across iterations with GPT-4 but degraded with weaker models like GPT-3.5 and Mixtral. Recursive structure alone is not enough. The base model must be capable enough to improve the mechanism. This implies that harness improvement enables better deployment of the model but intelligence is still the core.”

同一种自我改进方法,配 GPT-4 会越迭代越好,配 GPT-3.5 和 Mixtral 反而越迭代越差。所以基座模型必须「足够有能」,这套机制才起作用——框架能改进部署,但智能本身仍是核心。

4. 后训练耦合(posttraining coupling)

“Models get posttraining coupled to the harness they were trained against.”

模型会被「后训练」绑死在它训练时所用的那套框架上。

具体表现:“changing tool logic leads to worse model performance… A truly intelligent model should have little trouble switching between patch methods, but training with a harness in the loop creates this overfitting.”

2026 年这甚至成了公开产品实践——Meta 在 Muse Spark 1.2 的发布说明中写明: “We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together.”

Meta 公开写明他们把模型和自家编程工具一起联合训练,以保证二者配对使用时表现最好。这是「后训练耦合」最直白的一次公开产品实践。

5. 最有设计价值的一条:组件会过期

“Every component in a harness encodes an assumption about what the model can’t do on its own. When the model gets better at something, that component becomes load-bearing for nothing and should come out.”

这条可以直接迁移到”参数与能力”的讨论:能力维度会随模型代际变化,补偿机制必须跟着删。

6. 两个独立的轴(Lin et al. 2026,经 Lilian Weng 转述)

能力跨模型尺寸的表现
harness-updating(写出有用的框架改动)基本持平——Qwen3.5-9B 写的 skill 在程序结构上与 Opus 4.6 写的同构
harness-benefit(把改好的框架用起来)非单调,中间档模型受益最大;取决于”及时正确调用 skills/tools + 长程指令遵循”

这一条是”参数大不等于工具调用强”的机制解释:写框架靠的是结构化能力(小模型也有),用好框架靠的是调度与长程遵循(与参数不单调相关)。

7. 怎么做判断

  • 看到”某模型在某某智能体榜上 80%“,先问用的是什么 harness。
  • 选模型时,在你自己的框架里跑你自己的任务,别直接抄榜单。
  • 换框架后掉分,先怀疑耦合,再怀疑模型。
  • 升级模型后,回头检查框架里哪些补偿机制该删了。

相关页面

  • Harness工程 —— wiki 已有概念页
  • Agent —— wiki 已有概念页
  • 工具调用 —— 受分责影响最大的一维
  • Terminal-Bench与Harness分责 —— 原始榜单数据