Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
Zhongzhi Yu, Zheng Wang, Yonggan Fu, Huihong Shi, Khalid Shaikh, Yingyan Celine Lin
Abstract
Attention is a fundamental component behind the remarkable achievements of large language models (LLMs). However, our current understanding of the attention mechanism, especially regarding how attention distributions are established, remains limited. Inspired by recent studies that explore the presence of attention sink in the initial token, which receives disproportionately large attention scores despite their lack of semantic importance, this work delves deeper into this phenomenon. We aim to provide a more profound understanding of the existence of attention sinks within LLMs and to uncover ways to enhance the achievable accuracy of LLMs by directly optimizing the attention distributions, without the need for weight finetuning. Specifically, this work begins with comprehensive visualizations of the attention distributions in LLMs during inference across various inputs and tasks. Based on these visualizations, to the best of our knowledge, we are the first to discover that (1) attention sinks occur not only at the start of sequences but also within later tokens of the input, and ( 2 ) not all attention sinks have a positive impact on the achievable accuracy of LLMs. Building upon our findings, we propose a training-free Attention Calibration Technique (ACT) that automatically optimizes the attention distributions on the fly during inference in an input-adaptive manner. Extensive experiments validate that ACT consistently enhances the accuracy of various LLMs across different applications. Specifically, ACT achieves an average improvement of up to 7.30% in accuracy across different datasets when applied to Llama-30B. Our code is available at https: //github.com/GATECH-EIC/ACT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5799b427-4bfa-4d27-9293-a64ccefc0cd6Cited by top-tier papers38
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- What are you sinking? A geometric approach on attention sinkValeria Ruscio, Umberto Nanni, Fabrizio SilvestriNeurIPS 2025 · 29 citations
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language ModelsJiayun Luo, Wan-Cyuan Fan, Lyuyang Wang, Xiangteng He et al.ICLR 2026 · 19 citations
- Attention Sinks: A 'Catch, Tag, Release' Mechanism for EmbeddingsStephen Zhang, Mustafa Khan, Vardan PapyanNeurIPS 2025 · 18 citations
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without TrainingFeijiang Han, Xiaodong Yu, Jianheng Tang, Delip Rao et al.ICLR 2026 · 17 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMsXiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu et al.EMNLP 2025
- When Attention Sink Emerges in Language Models: An Empirical ViewXiangming Gu, Tianyu Pang, Chao Du, Qian Liu et al.ICLR 2025
- Identifying and Evaluating Inactive Heads in Pretrained LLMsPedro Sandoval-Segura, Xijun Wang, Ashwinee Panda, Micah Goldblum et al.ICLR 2026 · 8 citations
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language ModelsJakub Binkowski, Kamil Adamczewski, Tomasz KajdanowiczICML 2026
- ART: Attention Replacement Technique to Improve Factuality in LLMsZiqin Luo, Yihao Quan, Xiaofeng Zhang, Xiaosong Yuan et al.ACL 2026 · 1 citation
