Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
Zhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang, Hongming Zhang, Chenlong Deng, Shuaiyi Li, Dong Yu
Abstract
Large language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a013c1f3-635c-4c54-b25e-6c4174b4de29Cited by top-tier papers13
- KVLink: Accelerating Large Language Models via Efficient KV Cache ReuseJingbo Yang, Bairu Hou, Wei Wei, Yujia Bao et al.NeurIPS 2025 · 83 citations
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative PruningHanzhen Wang, Jiaming Xu, Yushun Xiang, Jiayi Pan et al.ICML 2026 · 32 citations
- MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-HeadKewei Zhang, Ye Huang, Yufan Deng, Jincheng Yu et al.ICLR 2026 · 6 citations
- SkyLadder: Better and Faster Pretraining via Context Window SchedulingTongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen et al.NeurIPS 2025 · 6 citations
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge et al.ACL 2026 · 4 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context ParallelismTao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun et al.ICLR 2026
- Parallel Context Windows for Large Language ModelsNir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram et al.ACL 2023 · 37 citations
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same CoinEnrique Queipo-de-Llano, Alvaro Arroyo, Federico Barbero, Xiaowen Dong et al.ICLR 2026 · 56 citations
- UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-TilingHaoyu Yang, Zan Zong, Yuyang Jin, Kinman Lei et al.SC 2025 · 1 citation
- Selective Attention Improves TransformerYaniv Leviathan, Matan Kalman, Yossi MatiasICLR 2025
