一致性指标:pass^k 与 pass@k
两个长得几乎一样、含义完全相反,而且哪个才对取决于你在做什么的指标。
1. 定义
| 指标 | 含义 | k 增大时 |
|---|---|---|
| pass@k(pass at k) | k 次尝试中至少一次成功 | 分数上升 |
| pass^k(pass hat k) | k 次尝试全部成功 | 分数下降 |
τ-bench 论文(arXiv:2406.12045)的原话:
“For tasks like code generation with good verification techniques (unit tests), the community has defined the pass@k metric as the chance that at least one out of k i.i.d. task trials is successful, which captures the trend of agents enabling discovery of solutions with scaling of inference-time compute. For real-world agent tasks requiring reliability and consistency like customer service, we propose a new metric – pass^k, defined as the chance that all k i.i.d. task trials are successful, averaged across tasks.”
2. Anthropic 官方的选择指南
“pass@k measures the likelihood that an agent gets at least one correct solution in k attempts. As k increases, pass@k score rises… pass^k measures the probability that all k trials succeed. As k increases, pass^k falls since demanding consistency across more trials is a harder bar to clear. If your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ≈ 42%. This metric especially matters for customer-facing agents where users expect reliable behavior every time.”
pass@k 随 k 增大而上升(多试几次总有一次中),pass^k 随 k 增大而下降(要求次次都对,次数越多越难)。举例:单次成功率 75%,跑 3 次全对的概率只有 0.75³ ≈ 42%。
“pass@k for tools where one success matters, pass^k for agents where consistency is essential.”
3. 为什么这个区别能改变结论
τ-bench 的实测(function calling 方法):
| 模型 | τ-retail pass^1 | 注 |
|---|---|---|
| gpt-4o | 61.2% | 看起来”过半了” |
| gpt-4o 的 pass^8 | < 25% | 连续 8 次都成功的概率不到四分之一 |
论文的三处表述:
摘要:“even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail)” 引言:“…to as low as ∼25% for pass^8 on τ-retail for the same model” §5.1:“Even for the best-performing gpt-4o function calling agent which has a >60% average task success, pass^8 drops to <25%”
直觉陷阱:单次成功率 61% 听起来是”三分之二的时候能用”;但如果你要把它接进客服系统,每 8 个用户里只有 2 个能得到全程正确的服务。
4. 一个进阶版本:双控环境(τ²-bench)
τ²-bench(arXiv:2506.07982)把环境从”只有智能体能动手”改成”智能体和用户都能动手改世界”:
“Existing benchmarks … simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate.”
结果:从 no-user 切到双控,gpt-4.1 掉 18 个百分点,o4-mini 掉 25 个百分点。
“Our findings reveal a substantial performance decrease (around 20% pass^1) when agents must shift from autonomous operation to guiding a user.”
把任务从「模型自己干」改成「模型要引导用户一起干」,单次成功率会掉大约 20 个百分点。
5. 与其他维度的关系
| 概念 | 关系 |
|---|---|
| 推理期缩放 | pass@k 正是”多试几次”这条路线的度量;它提升 pass@k,但不改善 pass^k |
| 工具调用 | 工具调用的错误会累积,所以长程智能体任务的 pass^k 衰减尤其严重 |
| 指令遵循 | 领域规则缺失导致的不稳定,直接体现在 pass^k 上 |
6. 使用建议
- 做评测/做选型时,先问自己:一次成功够吗?
- 写代码找解法 → pass@k
- 面向用户的服务、自动化流水线 → pass^k
- 报告跑分时务必注明是哪一个。这两个指标的分母相同、语义相反,混用会得出完全相反的结论。
- 至少要报 pass^1 和 pass^k 两个数,只看前者会系统性高估可靠性。