一致性指标:pass^k 与 pass@k

两个长得几乎一样、含义完全相反,而且哪个才对取决于你在做什么的指标。


1. 定义

指标含义k 增大时
pass@k(pass at k)k 次尝试中至少一次成功分数上升
pass^k(pass hat k)k 次尝试全部成功分数下降

τ-bench 论文(arXiv:2406.12045)的原话:

“For tasks like code generation with good verification techniques (unit tests), the community has defined the pass@k metric as the chance that at least one out of k i.i.d. task trials is successful, which captures the trend of agents enabling discovery of solutions with scaling of inference-time compute. For real-world agent tasks requiring reliability and consistency like customer service, we propose a new metric – pass^k, defined as the chance that all k i.i.d. task trials are successful, averaged across tasks.”

2. Anthropic 官方的选择指南

“pass@k measures the likelihood that an agent gets at least one correct solution in k attempts. As k increases, pass@k score rises… pass^k measures the probability that all k trials succeed. As k increases, pass^k falls since demanding consistency across more trials is a harder bar to clear. If your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ≈ 42%. This metric especially matters for customer-facing agents where users expect reliable behavior every time.”

pass@k 随 k 增大而上升(多试几次总有一次中),pass^k 随 k 增大而下降(要求次次都对,次数越多越难)。举例:单次成功率 75%,跑 3 次全对的概率只有 0.75³ ≈ 42%。

“pass@k for tools where one success matters, pass^k for agents where consistency is essential.”

3. 为什么这个区别能改变结论

τ-bench 的实测(function calling 方法):

模型τ-retail pass^1注
gpt-4o61.2%看起来”过半了”
gpt-4o 的 pass^8< 25%连续 8 次都成功的概率不到四分之一

论文的三处表述:

摘要:“even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail)” 引言:“…to as low as ∼25% for pass^8 on τ-retail for the same model” §5.1:“Even for the best-performing gpt-4o function calling agent which has a >60% average task success, pass^8 drops to <25%”

直觉陷阱:单次成功率 61% 听起来是”三分之二的时候能用”;但如果你要把它接进客服系统,每 8 个用户里只有 2 个能得到全程正确的服务。

4. 一个进阶版本:双控环境(τ²-bench)

τ²-bench(arXiv:2506.07982)把环境从”只有智能体能动手”改成”智能体和用户都能动手改世界”:

“Existing benchmarks … simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate.”

结果:从 no-user 切到双控,gpt-4.1 掉 18 个百分点,o4-mini 掉 25 个百分点。

“Our findings reveal a substantial performance decrease (around 20% pass^1) when agents must shift from autonomous operation to guiding a user.”

把任务从「模型自己干」改成「模型要引导用户一起干」,单次成功率会掉大约 20 个百分点。

5. 与其他维度的关系

概念关系
推理期缩放pass@k 正是”多试几次”这条路线的度量;它提升 pass@k,但不改善 pass^k
工具调用工具调用的错误会累积,所以长程智能体任务的 pass^k 衰减尤其严重
指令遵循领域规则缺失导致的不稳定,直接体现在 pass^k 上

6. 使用建议

  • 做评测/做选型时,先问自己:一次成功够吗?
    • 写代码找解法 → pass@k
    • 面向用户的服务、自动化流水线 → pass^k
  • 报告跑分时务必注明是哪一个。这两个指标的分母相同、语义相反,混用会得出完全相反的结论。
  • 至少要报 pass^1 和 pass^k 两个数,只看前者会系统性高估可靠性。

相关页面