自回归生成(Autoregressive Generation)
一句话定义:自回归生成(Autoregressive Generation)是大模型产出文本的方式——一次只生成一个词元(Token),且每一个词元都以前一个词元为条件,因此必须顺序进行。它强大、通用,但天生慢、天生贵,这也是 Jev 要绕开它的原因。
官方发布博客原话:“Sequential. Generates one token at a time, each conditioned on the last.”
中文:顺序式。一次生成一个 token,每个 token 都以前一个为条件。
1. 机制:为什么必须一步一步来
官方对这条路径的描述只有一句,但把关键全说了——每个 token 以前一个为条件(each conditioned on the last)。要产出 N 个 token,就必须走 N 个有先后依赖的步骤:不知道第 k 个 token,就无法算第 k+1 个。这是自回归的”自”——用自己上一步的输出当作下一步的输入。
Jev 走的是另一条路,官方把二者直接并置:
“Jev outputs all probabilities in parallel instead of autoregressively generating by token.”
一边是”逐 token 顺序生成”,一边是”全部概率并行输出”(见 并行采样)。
2. 为什么慢、为什么贵
慢。 官方给出的端到端响应时间是:
“End-to-end response time is 3 to 329 seconds for frontier models. Fast enough for interfacing with humans, but a big bottleneck when integrated in code.”
(前沿模型的端到端响应是 3 到 329 秒。对人交互够快,但集成进代码里就是大瓶颈。)对照 Jev:
“End-to-end response time is 70ms-500ms for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.”
| 大模型 | Jev | |
|---|---|---|
| 回答一次要多久 | 3~329 秒 | 70~500 毫秒(多数约 100 毫秒) |
| 生成的粒度 | 一个一个 token 往外写 | 所有答案一次算完 |
贵。 大模型按”输入的字 + 输出的字”双向收费,而且输出更贵:
“Cost — Input tokens: from 10 / MTok. Output tokens: ~5x more expensive than input tokens.”
(输入词元 10 / MTok;输出词元比输入贵约 5 倍。)Jev 只有输入收费:42 / 十亿词元),输出免费。官方称其输入价比 Claude Fable 5.1 低 238 倍。
首页那组工作流对照更直观:Jev 0.013880 / 8.566s(按这两个数相除,约贵 170 倍、慢 75 倍;官方对同一批 workflow 给出的综合口径是 “193.6x Faster, 444.6x Cheaper.”)。
3. 生成字符串的隐性代价
官方对”字符串”的定位很清醒——不是不好,而是贵且有风险:
“Strings are extremely powerful and general, but costly. ‘Giving up’ strings actually gives us a lot of superpowers!”
放弃生成能力,换来的是快、便宜、类型安全。对系统集成者来说,真正的痛点在”下游摩擦”:
“Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.”
(字符串灵活到可以是任何东西:聊天回复、代码、幻觉、拒答,甚至类型安全的结构化值。要被软件使用,回复必须先解析再校验;而且总存在”跑偏”的风险。)这就是为什么会有”让模型按 JSON 格式输出”的纠结——官方在 Introduction 里称之为一种错配:
“When you need a model to make a judgment that your code will consume, that creates a mismatch: you are coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.”
(你需要模型做一个”代码要消费”的判断时,就产生了错配:你在强迫一个文本生成系统输出结构化决策,再把结果解析回代码能依赖的东西。)
4. 多问几个问题,生成模型要付出什么
对自回归模型,“多问一个问题”往往意味着要么更长的一次生成、要么更多轮往返,两者都要串行地花时间与输出词元。而 Jev 因为并行采样,“Adding questions barely changes the response time… Asking a question you might not need is close to free.”官方 cookbook 实测:13 个问题一次问 vs 拆 13 次问,后者要 2.71 秒、0.000497(便宜 12.2 倍、快 10.0 倍)。
这也是官方”快慢分工”论点的由来:让生成模型只负责它不可替代的”写”,把大量判断交给判别式模型(见 判别式模型)。
5. 与 Jev 的关系:分工,而不是取代
Jev 明确不做生成,官方在缺陷清单里单列一条并给出替代方案:
“
jev-1.13is not trained to generate text. While you can force it to by chaining choices, this will not work well and will be very slow.”Instead: when the answer space is bounded, turn extraction into a Choice over the options rather than asking for the value itself. If you really need to generate text… there are other models for that.
值得注意:强行让 Jev”逐 Choice 串起来当生成器用”——本质就是在判别模型上模拟一次自回归生成——结果依然是又慢又差。这反过来印证了自回归的成本来自机制本身,而不是某个模型实现得不好。
正确的用法是混搭。官方智能家居 demo 就是范例:
“When TypeSafe determines that the user query is a request for general information or conversation, the system calls an LLM to generate a freeform response… The initial TypeSafe response is so fast compared to the LLM response that it adds negligible latency to the overall system.”
(当 TypeSafe 判定用户是在问信息或闲聊时,系统才调用 LLM 去生成自由文本回复;TypeSafe 那一步相对 LLM 快到可以忽略,几乎不给整体延迟加负担。)分工原则:判别模型做过滤、路由、把关、打分;生成模型做写作。
相关来源
- 官方发布博客《Introducing System One Models & Jev》(Sampling / Cost / Speed 对照表):
processed/jev-原始资料/官方-网页/o04-blog-sysone.txt(https://typesafe.ai/blog/system-one) - 官网首页(193.6x / 444.6x 与 0.114s vs 8.566s):
processed/jev-原始资料/官方-网页/o01-home.txt(https://typesafe.ai) - 官方文档《Introduction》(“mismatch” 段):
processed/jev-原始资料/官方-文档/introduction.md(https://docs.typesafe.ai/introduction) - 官方文档《AI primer》(RLHF / RLVR / RLCD):
processed/jev-原始资料/官方-文档/introduction__machine-learning-primer.md - 官方文档《Primitives (Questions)》:
processed/jev-原始资料/官方-文档/primitives.md - 官方文档《Jev 1.13 jaggedness · Generation》:
processed/jev-原始资料/官方-文档/model-jaggedness__jev-1.13.md - 官方文档《Smart home assistant demo》:
processed/jev-原始资料/官方-文档/demos__smart-home.md