Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals
Roy Xie, Junlin Wang, Paul Rosu, Chunyuan Deng, Bolun Sun, Zihao Lin, Bhuwan Dhingra
Abstract
Large language models (LLMs) process entire input contexts indiscriminately, which is inefficient when the information required to answer a query is localized within the context. We present dynamic context cutoff, a novel method enabling LLMs to self-terminate processing upon acquiring sufficient task-relevant information. Through analysis of model internals, we discover that specific attention heads inherently encode "sufficiency signals" -detectable through lightweight classifiers -that predict when critical information has been processed. This reveals a new efficiency paradigm: models' internal understanding naturally dictates processing needs rather than external compression heuristics. Comprehensive experiments across six QA datasets (up to 40K tokens) with three model families (LLaMA/Qwen/Mistral, 1B-70B) demonstrate 3.4% accuracy improvement while achieving 1.33× token reduction on average. Furthermore, our method demonstrates superior performance compared to other context efficiency methods at equivalent token reduction rates. Additionally, we observe an emergent scaling phenomenon: while smaller models require probing for sufficiency detection, larger models exhibit intrinsic self-assessment capabilities through prompting. Code is available at https://github.com/ruoyuxie/when-to-stop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe8e78fa-5114-4c64-9e55-7ca7da3a6a98Builds on10
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
Related papers
- DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question AnsweringJiakai Li, Rongzheng Wang, Yizhuo Ma, Shuang Liang et al.NeurIPS 2025 · 8 citations
- LLoCO: Learning Long Contexts OfflineSijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu et al.EMNLP 2024 · 3 citations
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye et al.NeurIPS 2025 · 16 citations
- Dodo: Dynamic Contextual Compression for Decoder-only LMsGuanghui Qin, Corby Rosset, Ethan C. Chau, Nikhil Rao et al.ACL 2024
