一句话摘要:NoLiMa 把”大海捞针”里的字面匹配漏洞堵上,结果 13 个标称 ≥128K 的模型中,11 个在 32K(仅标称的四分之一)就掉到短上下文表现的一半以下;连最强的 GPT-4o 也从 99.3% 掉到 69.7%。

原始标题

NoLiMa: Long-Context Evaluation Beyond Literal Matching(LMU Munich / Adobe,arXiv:2502.05167,ICML 2025)

来源:https://arxiv.org/abs/2502.05167


它解决什么问题

经典大海捞针(Needle In A Haystack, NIAH)测试有个致命漏洞:针和草堆之间往往存在字面重叠,模型可以靠关键词匹配”作弊”,不需要真正理解。NoLiMa 要堵上这个漏洞。

关键要点

1. 对 NIAH 的批评

“in these benchmarks, models can exploit existing literal matches between the needle and haystack to simplify the task. To address this, we introduce NoLiMa, a benchmark extending NIAH with a carefully designed needle set, where questions and needles have minimal lexical overlap, requiring models to infer latent associations to locate the needle within the haystack.”

举例:问题问”哪本书的作者去过东京”,而针是”《XX》的作者曾在东京生活”——“东京”这个词只出现在问题里,不出现在针里,模型必须推断出关联。

2. 核心数字

“We evaluate 13 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 11 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%.”

我们测了 13 个自称支持 128K 以上上下文的模型。它们在短上下文(<1K)下表现很好,但随着长度增加会显著退化。在 32K 时,11 个模型已经掉到自己短上下文成绩的 50% 以下;连表现最好的 GPT-4o 也从几乎满分掉到了 69.7%。

位置表现
< 1K(短上下文)强基线(GPT-4o 为 99.3%)
32K(标称的 1/4)13 个模型中 11 个掉到基线 50% 以下;GPT-4o 掉到 69.7%

3. 为什么这条数据重要

它直接戳破了”上下文窗口 = 可用长度”的等式。标称窗口是”能塞多少”,不是”能用到多少”。

重要引用(英文原文)

“At 32K, for instance, 11 models drop below 50% of their strong short-length baselines.”

比如在 32K 这个长度(只有标称长度的四分之一),13 个模型里有 11 个掉到了自己短上下文基线的一半以下。

交叉验证:RULER(arXiv 2404.06654,NVIDIA,COLM 2024)

“The needle-in-a-haystack (NIAH) test … this simple retrieval-based test is indicative of only a superficial form of long-context understanding. … We evaluate 17 long-context LMs with 13 representative tasks in RULER. Despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases. While these models all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K.”

RULER 在 vanilla NIAH 之外新增 multi-hop tracing(多跳追踪) 与 aggregation(聚合) 两类任务,测的是”超出单纯检索”的行为。

边界与争议

  • “有效上下文 = 标称的 50–65%“这类逐模型百分比表在本次抓取中只见于二手聚合站(techjacksolutions、elvex),两源口径不同且无一手 RULER 报告支撑。只可引用定性结论与本页的一手数字。
  • 这是 2025 年初的测试(GPT-4o 时代)。2026 年的模型在长上下文上已有实质改进,但”标称 ≠ 有效”的结构性差距依然存在。

与本文其他页面的关系