Attention Basin: Why Contextual Position Matters in Large Language Models
Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu, Wei Liu, Jian Luan, Wanxia Cao, Ying Shen
Abstract
The performance of Large Language Models (LLMs) is significantly sensitive to the contextual position of information in the input. To investigate the mechanism behind this positional bias, our extensive experiments reveal a consistent phenomenon we term the attention basin: when presented with a sequence of structured items (e.g., retrieved documents or few-shot examples), models systematically assign higher attention to the items at the beginning and end of the sequence, while neglecting those in the middle. Crucially, our analysis further reveals that allocating higher attention to critical information is key to enhancing model performance. Based on these insights, we introduce Attention-Driven Reranking (AttnRank), a two-stage framework that (i) estimates a model's intrinsic positional attention preferences using a small calibration set, and (ii) reorders retrieved documents or few-shot examples to align the most salient content with these high-attention positions. AttnRank is a model-agnostic, training-free, and plug-and-play method with minimal computational overhead. Experiments on multi-hop QA and few-shot in-context learning tasks demonstrate that AttnRank achieves substantial improvements across 10 large language models of varying architectures and scales, without modifying model parameters or training procedures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0291499d-6503-4e59-9ea3-7c6f6ab1e2e6Cited by top-tier papers2
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning AbilitiesJiayi Kuang, Haojing Huang, Yinghui Li, Xinnian Liang et al.NeurIPS 2025 · 11 citations
- Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition BottleneckMeiru Zhang, Zaiqiao Meng, Nigel CollierACL 2026 · 1 citation
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
Related papers
- Attention in Large Language Models Yields Efficient Zero-Shot Re-RankersShijie Chen, Bernal Jimenez Gutierrez, Yu SuICLR 2025
- Where to show Demos in Your Prompt: A Positional Bias of In-Context LearningKwesi A. Cobbina, Tianyi ZhouEMNLP 2025
- LongRanker: Efficient One-Pass Document Reranking with Long-Context Large Language ModelsChangjiang Zhou, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.WWW 2026
- Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric TasksMohsen Dehghankar, Abolfazl AsudehKDD 2025 · 1 citation
- Few-shot Reranking for Multi-hop QA via Language Model PromptingMuhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee et al.ACL 2023 · 5 citations
