知识容量(Knowledge Capacity)

一句话定义:一个模型最多能装下多少比特的事实知识。知识容量缩放定律论文 用受控合成数据测出的答案是:每参数 2 bits,且这是上限。


1. 核心结论:2 bits/parameter

“Through multiple controlled datasets, we establish that language models can and only can store 2 bits of knowledge per parameter, even when quantized to int8… Consequently, a 7B model can store 14B bits of knowledge, surpassing the English Wikipedia and textbooks combined based on our estimation.”

计量单位是 (实体, 属性, 值) 三元组,例如 (USA, capital, Washington D.C.)。

“only can(只能)“三个字很关键:这是上限,不是平均值。训练不足时会显著低于 2 bits。

2. 影响容量的五个因素

论文摘要明确列出:

  1. 训练时长(training duration)——知识要反复曝光才能固化
  2. 模型架构(model architecture)
  3. 量化(quantization)——int8 下 2 bits/param 仍成立
  4. 稀疏约束(sparsity constraints such as MoE)
  5. 数据信噪比(data signal-to-noise ratio)

3. 两个反直觉的发现

架构上,新旧未必是优劣:

“The GPT-2 architecture, with rotary embedding, matches or even surpasses LLaMA/Mistral architectures in knowledge storage, particularly over shorter training durations. This arises because LLaMA/Mistral uses GatedMLP, which is less stable and harder to train.”

带旋转位置编码的 GPT-2 架构,在知识存储上匹配甚至超过 LLaMA/Mistral 架构,训练时长较短时尤其明显。原因是 LLaMA/Mistral 用的 GatedMLP 更不稳定、更难训练。这条很反直觉——「更先进的架构」反而更不擅长存知识。

数据组织方式本身就是杠杆:

“Prepending training data with domain names (e.g., wikipedia.org) significantly increases a model’s knowledge capacity. Language models can autonomously identify and prioritize domains rich in knowledge, optimizing their storage capacity.”

工程含义:同样的数据,光是标注来源就能提高容量。这意味着数据工程的价值独立于参数量。

4. 装得下 ≠ 用得上

这是理解”能力维度”的关键分野。知识操纵与思维链论文 证明:

  • 模型能把每个人的出生月份答对 100%、训了 25,000 条样本,却回答不了”这个人出生在偶数月吗”——除非先让它说出月份再判断。
  • “Improving model’s knowledge extraction don’t improve its manipulation ability.”
  • 反向检索(“谁的 X 属性等于 T?“)无论怎么训练都近乎 0%:“language models cannot be used as databases.”

所以正确的心智模型是:知识容量是”仓库面积”,推理操纵是”分拣能力”——两者独立,前者不带来后者。

5. 对 MoE 的特别提醒

稀疏约束(MoE)被明确列为影响容量的因素之一,因此:

“总参数 × 2 bits”不能直接套用到 MoE 模型上。671B 总参 / 37B 激活的 DeepSeek-V3,其知识容量既不等于 671B 稠密模型,也不等于 37B 稠密模型——具体数值在该论文的 12 条结论中,本次只核实到摘要层面。

6. 常见误用

  • ❌ “7B 模型 = 14B bits = 能背下整个维基百科,所以知识问题都解决了”——装得下不等于查得到,见第 4 节。
  • ❌ “量化会按比例砍掉知识”——int8 量化不损害知识容量(论文一手结论)。
  • ❌ “换个更大参数就能解决知识错误”——知识错误更多来自训练数据的信噪比和提取路径,不是容量不够。

相关页面