NoLiMa: Long-Context Evaluation Beyond Literal Matching
Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui, Ryan A. Rossi, Seunghyun Yoon, Hinrich Schütze
摘要
Recent large language models (LLMs) support long contexts ranging from 128K to 1M tokens. A popular method for evaluating these capabilities is the needle-in-a-haystack (NIAH) test, which involves retrieving a "needle" (relevant information) from a "haystack" (long irrelevant context). Extensions of this approach include increasing distractors, fact chaining, and in-context reasoning. However, in these benchmarks, models can exploit existing literal matches between the needle and haystack to simplify the task. To address this, we introduce NOLIMA, a benchmark extending NIAH with a carefully designed needle set, where questions and needles have minimal lexical overlap, requiring models to infer latent associations to locate the needle within the haystack. We evaluate 13 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 11 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almostperfect baseline of 99.3% to 69.7%. Our analysis suggests these declines stem from the increased difficulty the attention mechanism faces in longer contexts when literal matches are absent, making it harder to retrieve relevant information. Even models enhanced with reasoning capabilities or CoT prompting struggle to maintain performance in long contexts. We publicly release the dataset and evaluation code at https://github.com/adoberesearch/NoLiMa . 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code UnderstandingAdam Storek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava 等ACL 2026 · 被引用 4 次
- AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMsLara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas CavigelliNeurIPS 2025 · 被引用 3 次
- CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor DensityDaniel Kaiser, Arnoldo Frigessi, Ali Ramezani-Kebrya, Benjamin RicaudICLR 2026 · 被引用 3 次
- A Multi-Agent Framework for High-Interaction Terminal SimulationKai Wei, Yuwen Cui, Kehan Shen, Hua Wei 等ACL 2026
- L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information ScalingZhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo 等ICML 2026
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet 等NeurIPS 2024 · 被引用 271 次
- Random-Access Infinite Context Length for TransformersAmirkeivan Mohtashami, Martin JaggiNeurIPS 2023 · 被引用 207 次
相关 Paper
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long ContextsYifei Yu, Qian-Wen Zhang, Lingfeng Qiao, Di Yin 等EMNLP 2025 · 被引用 2 次
- HELMET: How to Evaluate Long-context Models Effectively and ThoroughlyHoward Yen, Tianyu Gao, Minmin Hou, Ke Ding 等ICLR 2025
- One Thousand and One Pairs: A "novel" challenge for long-context language modelsMarzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal 等EMNLP 2024 · 被引用 6 次
- NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured TablesLanrui Wang, Mingyu Zheng, Hongyin Tang, Zheng Lin 等NeurIPS 2025 · 被引用 16 次
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
