一句话摘要:Chroma 用 18 个模型、8 种输入长度、11 个答案位置做系统实验,结论是没有例外——模型不会均匀使用上下文,输入越长表现越不可靠;而且反直觉的是,逻辑连贯的文本比打乱的文本更难。
原始标题
Context Rot: How Increasing Input Tokens Impacts LLM Performance(Kelly Hong, Anton Troynikov, Jeff Huber, Chroma 技术报告,2025-07-14)
来源:https://research.trychroma.com/context-rot
它解决什么问题
厂商在宣传”1M 上下文窗口”。但”能塞进去 1M token”和”塞进去之后还能正常思考”是两件事。Chroma 要量化后者。
关键要点
1. 实验规模与被测模型
“In this report, we evaluate 18 LLMs, including the state-of-the-art GPT-4.1, Claude 4, Gemini 2.5, and Qwen3 models.”
在这份报告里,我们评测了 18 个大语言模型,包括当时最先进的 GPT-4.1、Claude 4、Gemini 2.5 和 Qwen3。
被测清单(附录):Claude Opus 4 / Sonnet 4 / Sonnet 3.7 / Sonnet 3.5 / Haiku 3.5;o3 / GPT-4.1 / GPT-4.1 mini / GPT-4.1 nano / GPT-4o / GPT-4 Turbo / GPT-3.5 Turbo;Gemini 2.5 Pro / 2.5 Flash / 2.0 Flash;Qwen3-235B-A22B / Qwen3-32B / Qwen3-8B。
“For every unique combination of needle type, haystack topic, and haystack structure, we test each model across: 8 input lengths / 11 needle positions.”
每一种「针的类型 × 草堆主题 × 草堆结构」的组合,都在 8 种输入长度和 11 个插入位置上各测了一遍——所以这个结论不是单一设置下的偶然结果。
2. 总发现
“Our results reveal that models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.”
模型不是均匀地使用它全部的上下文窗口;随着输入变长,它的表现会越来越不可靠。
3. 大海捞针(NIAH)为什么不够用
“NIAH is fundamentally a simple retrieval task … While scalable, this benchmark typically assesses direct lexical matching, which may not be representative of flexible, semantically oriented tasks.”
大海捞针(NIAH)本质上是个简单的检索任务。……它易于规模化,但考的通常是字面上的直接匹配,不能代表需要语义理解的灵活任务。
4. 五项发现(逐条)
- 所有实验中,表现随输入长度增加而持续下降。
- needle 与问题的相似度越低,下降越快。
- 干扰项(distractor)的影响不均匀,且随长度增加更明显:
“Even a single distractor reduces performance relative to the baseline (needle only), and adding four distractors compounds this degradation further.”
哪怕只加一条干扰信息,表现就会比「只有针」的基线下降;而加到四条时,这种退化会进一步叠加放大。
- needle-haystack 相似度的影响不是统一的。
- haystack 的结构会稳定影响模型如何处理长输入。
5. 模型家族之间的行为分化
“Claude models consistently exhibit the lowest hallucination rates. Specifically, Claude Sonnet 4 and Opus 4 are particularly conservative and tend to abstain when uncertain, explicitly stating that no answer can be found. In contrast, GPT models show the highest rates of hallucination, often generating confident but incorrect responses when distractors are present.”
这条对”能力维度”有直接含义:拒答能力(abstention)本身是一个独立维度,幻觉率低的模型不是”懂得多”,而是”更愿意说不知道”。
6. 反直觉发现:连贯文本比打乱文本更差
“Surprisingly, we find that structural coherence consistently hurts model performance.”
令人意外的是,文本的连贯结构会持续损害模型表现。这和「读得越通顺越容易理解」的直觉相反。
“Although it seems counterintuitive, models perform worse when the haystack preserves a logical flow of ideas. Shuffling the haystack and removing local coherence consistently improves performance.”
虽然反直觉,但草堆文本保持逻辑连贯时模型表现更差;把文本打乱、去掉局部连贯,表现反而稳定提升。
7. 真实长对话场景(LongMemEval)
“We use LongMemEval_s and filter for tasks … to end up with 306 total prompts. These prompts average out to ~113k tokens.” 而”Focused prompts average to ~300 tokens”。
他们从 LongMemEval_s 里筛出 306 条提示,平均长度约 11.3 万 token;对应的「聚焦版」提示平均只有约 300 token。
“We verify that the models are highly capable of succeeding on the focused inputs, then observe consistent performance degradation with the full inputs. … adding irrelevant context, and thereby adding an additional step of retrieval, significantly impacts a model’s ability to maintain reliable performance.”
翻译成工程语言:同一条信息,先检索出来再喂给模型,比整包丢进去效果好得多——这正是上下文工程(context engineering)的价值所在。
8. 报告的结论段
“Our results highlight the need for more rigorous long-context evaluation beyond current benchmarks, as well as the importance of context engineering. Whether relevant information is present in a model’s context is not all that matters; what matters more is how that information is presented.”
我们的结果说明「光有现有基准还不够,需要更严格的长上下文评测」;同时也说明了上下文工程的重要性——相关信息在不在上下文里还不是全部,更重要的是它怎么呈现。
重要引用(英文原文)
“Through our experiments, we demonstrate that LLMs do not maintain consistent performance across input lengths. Even on tasks as simple as non-lexical retrieval or text replication, we see increasing non-uniformity in performance as input length grows.”
通过实验我们证明,大语言模型在不同输入长度下的表现并不一致。哪怕是「非字面检索」或「照抄一段文本」这么简单的任务,随着输入变长,表现的不均匀性也会越来越明显。
边界与争议
- 这是厂商(向量数据库公司 Chroma)发布的技术报告,不是同行评议论文。但实验规模(18 模型 × 8 长度 × 11 位置)与公开方法使其被广泛引用。
- 它测的是 2025 年中期的模型,绝对数字会变;“随长度衰减”这个定性结论在 NoLiMa、RULER 等独立研究中得到交叉验证。
- 报告用的是 GPT-4.1 作为 LLM 裁判,作者声明与人类判断一致性 >99%。