投机式扇出(Speculative Fan-out)

一句话定义:把系统可能用到的所有问题——包括那些只在部分输入下才有意义的”投机性问题”——一次全部问出去,事后再用代码筛掉不相关的答案。 扇出(Fan-out)是”一次请求摊开问很多”,投机(Speculative)是”明知可能用不上也先问”。


1. 官方建议的原文

“Because TypeSafe supports sending many questions in a single API call, we recommend putting all of the questions your system needs in a single request, and then using code to decide what is relevant after the fact. All questions are evaluated in parallel, so adding more questions usually has little effect on response time.”

翻译:因为 TypeSafe 支持在一次 API 调用里发很多问题,我们建议把你系统需要的所有问题都放进同一个请求,事后用代码决定哪些是相关的。所有问题是并行(Parallel)评估的,所以多加几个问题通常几乎不影响响应时间。

patterns.md 的模式表把它概括为:“Send many questions in a single call, including speculative ones, and let your code decide what’s relevant”(一次调用发很多问题,包括投机性的,让代码决定哪些相关),收益是成本(Cost)、速度(Speed)。

2. 为什么”加问题几乎不增加耗时”

关键在于问题的评估方式:

  • 所有问题看到同一份状态(State),各自独立评估;
  • 一个问题的答案不会变成另一个问题的隐藏上下文——你可以增删问题而不改变其他问题的结果;
  • 因此新增问题的边际成本,只剩”多出来那几个问题的词元(Token)“,而输入词元本身很便宜。

primitives.md 的表述是:“Adding questions barely changes the response time and costs only the tokens for the extra questions, which are cheap. Asking a question you might not need is close to free.”(加问题几乎不改变响应时间,只多花那几个问题的词元,而它们很便宜。问一个你可能不需要的问题,代价接近于零。)

官方给出过量化的对照(cookbooks__parallel_questions.md,文档为英文维基百科 GDPR 条目,约 54,000 字符,13 个问题:8 个是非题 + 2 个选择题 + 3 个打分题):

策略调用次数成本总耗时
一次调用装 13 个问题1$0.0004970.27s
13 次调用、每次 1 个问题13$0.0060902.71s

结果:“batching: 12.2x cheaper, 10.0x faster”(合并成一次调用:便宜 12.2 倍、快 10.0 倍)。

同一个 cookbook 还验证了答案不会因为合并而改变:13 个问题各重复 5 次,两种策略下大多数答案逐次完全一致(标准差恰为 0.0);带采样噪声的两个问题(breach_72h、criminal_penalties)在两种策略下噪声量级相同。结论原文:“The noise is a property of the question, not of how you batch: batching neither shifts the answer nor adds variance.”(噪声是问题本身的属性,不是你如何分批造成的:分批既不移动答案,也不增加方差。)

3. 什么算”投机性问题”

patterns__fan-out.md 的支持工单分诊(Support Ticket Triage)示例:一次问 5 个问题——category(选择题)、bug_severity(打分题)、has_reproducible_steps(是非题)、refund_requested(是非题)、frustration(打分题)。官方原注:

“Speculative questions: bug_severity and has_reproducible_steps only matter if the ticket is a bug report. refund_requested only matters for billing. We include all upfront because additional questions usually have little effect on response time. If the ticket turns out to be a feature request, the bug severity result will be irrelevant, in which case your code path simply ignores it.”

翻译:bug_severity 和 has_reproducible_steps 只在工单是缺陷报告(bug report)时才有意义,refund_requested 只在账单类(billing)时才有意义。我们全都提前问了,因为额外问题通常几乎不影响响应时间。如果工单最后是功能请求,缺陷严重度的结果就不相关,代码路径直接忽略它就好。

代码侧做的就是”按分类结果筛”:

if category.choice == "bug_report":
    if bug_severity.score > 1.5 and bug_repro.noul > 0.6:
        escalate_to_engineering(ticket_id, severity="high")
    else:
        add_to_bug_backlog(ticket_id)
elif category.choice == "billing":
    if refund.noul > 0.7:
        route_to_billing_with_flag(ticket_id, refund_likely=True)
elif category.choice == "feature_request":
    log_feature_request(ticket_id)

官方的收尾是:“Everything needed for the full decision tree comes from one call. Speculative questions are ignored when irrelevant and save a round trip when they are not.”(整棵决策树需要的一切都来自同一次调用。 投机性问题在不相关时被忽略,在相关时省掉一次往返。)

4. 官方智能家居示例

demos__smart-home.md 说这个演示的主打模式就是投机式扇出:“Each user request is evaluated against a long list of questions, including many that will end up irrelevant for most requests.”(每条用户请求都会被拿去对照一长串问题,其中很多对多数请求最终都是无关的。)

对”Turn off all of the lights in the house”(关掉全屋的灯)这条请求,代码实际只需要 4 个答案:请求类别(智能家居指令)、目标域(全屋)、设备类型(灯)、对灯执行的动作(关闭)。

“Notice that the last question is written with the assumption that the user is issuing a command to lights, and we ask it before we know what the user is actually requesting. This is what we call a “speculative question” - we ask it before we even know if it’s relevant, allowing us to evaluate all questions in parallel and rely on code to filter out the irrelevant results after the fact. This is a key pattern for building systems that can handle a wide variety of user requests with a single set of questions.”

翻译:注意最后一个问题是假设用户在下达”对灯”的指令的前提下写的,而我们在还不知道用户到底在请求什么的时候就已经问了它。这就是我们说的”投机性问题”——在还不知道它相不相关时就先问,从而让所有问题并行评估,再靠代码事后过滤掉无关结果。这是用一组问题应对千变万化请求的关键模式。

演示里还明确批评了串行做法:

“The wrong way to do this would be to separate the questions in to multiple API calls, waiting to ask questions only once you are certain you need the answer… This approach optimizes for a minimum number of questions, but it ends up being much slower and more expensive than batching all of the questions in to one upfront API call.”

翻译:错误做法是把问题拆到多次 API 调用里,等确定需要答案了才去问。这种做法优化的是”问题的数量最少”,但它比把所有问题一次批量问出去要慢得多、贵得多。

5. 什么时候确实需要第二次请求

投机式扇出的前提是”问题彼此独立”。primitives.md 明确指出例外:如果后一个判断依赖前一个答案(要靠它去取更多数据、要决定状态由什么构成、要决定下一个问题的选项),那就必须在代码里发第二次请求。

“Two requests are the exception, not the rule. If the second request’s questions could have been asked against the original state, ask them in the first request and let the code ignore the ones it doesn’t need.”

翻译:两次请求是例外,不是常态。 如果第二个请求的问题本来就能对着原始状态问,那就放进第一个请求,让代码忽略不需要的那些。

官方举的三个真正需要第二次请求的例子:技能推荐(先对 182 个技能排序,再取前 3 个的正文重新判断)、结构恢复(先判断换行是否切断了句子,再合并成块并分类)、层级分类(用上一层选择题的答案决定下一次请求给哪些选项)。

6. 工程要点

  • 一次请求 = 一份状态 + 所有可能用得到的问题,别按”分支”拆调用。
  • 投机性问题的答案必须能被代码安全忽略——写代码时假设”它可能毫无意义”。
  • 问题要彼此独立:不要让一个问题偷偷依赖另一个的答案;真有依赖才发第二次请求。
  • 状态本身仍要精简:扇出的是问题,不是材料。材料里塞无关内容会拉低准确率(见 上下文腐坏)。
  • 官方提醒:编程智能体(Coding Agent)最容易犯”一次只问一个问题”的毛病,官方专门写了助手技能包来纠正这个习惯。

相关来源

  • processed/jev-原始资料/官方-文档/patterns__fan-out.md(官方模式定义、工单分诊示例)
  • processed/jev-原始资料/官方-文档/demos__smart-home.md(官方智能家居示例与”错误做法”对照)
  • processed/jev-原始资料/官方-文档/cookbooks__parallel_questions.md(13 问题、12.2x / 10.0x 实测)
  • processed/jev-原始资料/官方-文档/primitives.md(并行评估、何时需要第二次请求)
  • wiki/synthesis/Jev 知识.md(面向阅读的综合讲解)