Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?
Jonathan Roberts, Kai Han, Samuel Albanie
摘要
As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. In many real-world tasks, decisions depend on details scattered across collections of often disparate documents containing mostly irrelevant information. Long-context LLMs appear well-suited to this form of complex information retrieval and reasoning, which has traditionally proven costly and time-consuming. However, although the development of longer context models has seen rapid gains in recent years, our understanding of how effectively LLMs use their context has not kept pace. To address this, we conduct a set of retrieval experiments designed to evaluate the capabilities of 17 leading LLMs, such as their ability to follow threads of information through the context window. Strikingly, we find that many models are remarkably threadsafe: capable of simultaneously following multiple threads without significant loss in performance. Still, for many models, we find the effective context limit is significantly shorter than the supported context length, with accuracy decreasing as the context window grows. Our study also highlights the important point that token counts from different tokenizers should not be directly compared-they often correspond to substantially different numbers of written characters. We release our code and long context experimental data. INTRODUCTION 0 1 2 Context Length (million LLaMA 3.1 tokens)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language ModelsJunjie Wu, Gefei Gu, Yanan Zheng, Dit-Yan Yeung 等ACL 2025 · 被引用 4 次
- Relational Deep Dive: Error-Aware Queries Over Unstructured DataDaren Chao, Kaiwen Chen, Naiqing Guan, Nick KoudasVLDB 2026 · 被引用 3 次
它引用的顶会 Paper11
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet 等NeurIPS 2024 · 被引用 271 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 等NeurIPS 2024 · 被引用 212 次
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee 等ICLR 2024 · 被引用 131 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
相关 Paper
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu 等ACL 2024
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMsSangyeop Kim, Yohan Lee, Yongwoo Song, Kimin LeeACL 2025
- What is Wrong with Perplexity for Long-context Language Modeling?Lizhe Fang, Yifei Wang, Zhaoyang Liu, Chenheng Zhang 等ICLR 2025 · 被引用 2 次
- LongSafety: Evaluating Long-Context Safety of Large Language ModelsYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui 等ACL 2025 · 被引用 6 次
- How to Train Long-Context Language Models (Effectively)Tianyu Gao, Alexander Wettig, Howard Yen, Danqi ChenACL 2025
