Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
Yuhan Chen, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang, Yongbin Li, Rui Yan
摘要
In this paper, we demonstrate that an inherent waveform pattern in the attention allocation of large language models (LLMs) significantly affects their performance in tasks demanding a high degree of context awareness, such as utilizing LLMs for tool-use. Specifically, the crucial information in the context will be potentially overlooked by model when it is positioned in the trough zone of the attention waveform, leading to decreased performance. To address this issue, we propose a novel inference method named Attention Buckets. It allows LLMs to process their input through multiple parallel processes. Each process utilizes a distinct base angle for the rotary position embedding, thereby creating a unique attention waveform. By compensating an attention trough of a particular process with an attention peak of another process, our approach enhances LLM's awareness to various contextual positions, thus mitigating their risk of overlooking crucial information. In the largest tool-use benchmark, our method elevates a 7B model to achieve state-of-the-art performance comparable to that of GPT-4. On other benchmarks and some RAG tasks, which also demand a thorough understanding of contextual content, Attention Buckets also exhibited notable enhancements in performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional EncodingZhenyu Zhang, Runjin Chen, Shiwei Liu, Zhewei Yao 等NeurIPS 2024 · 被引用 99 次
- Attributive Reasoning for Hallucination Diagnosis of Large Language ModelsYuyan Chen, Zehao Li, Shuangjie You, Zhengyu Chen 等AAAI 2025 · 被引用 33 次
- Mixture of In-Context Experts Enhance LLMs' Long Context AwarenessHongzhan Lin, Ang Lv, Yuhan Chen, Chen Zhu 等NeurIPS 2024 · 被引用 25 次
- POSITION BIAS MITIGATES POSITION BIAS: Mitigate Position Bias Through Inter-Position Knowledge DistillationYifei Wang, Feng Xiong, Yong Wang, Linjing Li 等EMNLP 2025 · 被引用 16 次
- StreamingDialogue: Prolonged Dialogue Learning via Long Context Compression with Minimal LossesJianan Li, Quan Tu, Cunli Mao, Zhengtao Yu 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
相关 Paper
- Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention CalibrationZhongzhi Yu, Zheng Wang, Yonggan Fu, Huihong Shi 等ICML 2024 · 被引用 63 次
- LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional EncodingHaocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat 等ICML 2026
- Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingSenyu Han, Hongchuan Zeng, Kai Yu, Lu ChenICML 2025
- Attention Basin: Why Contextual Position Matters in Large Language ModelsZihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo 等ACL 2026 · 被引用 2 次
- Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsZhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang 等ACL 2025 · 被引用 22 次
