DAPE V2: Process Attention Score as Feature Map for Length Extrapolation
Chuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong, Jiankai Sun, Jingyao Li, Minbin Huang, Xiaozhe Ren, Michael Ng, Xin Jiang, Zhenguo Li, Yu Li
摘要
The attention mechanism is a fundamental component of the Transformer model, contributing to interactions among distinct tokens, in contrast to earlier feed-forward neural networks. In general, the attention scores are determined simply by the key-query products. However, this work's occasional trial (combining DAPE and NoPE) of including additional MLPs on attention scores without position encoding indicates that the classical key-query multiplication may limit the performance of Transformers. In this work, we conceptualize attention as a feature map and apply the convolution operator (for neighboring attention scores across different heads) to mimic the processing methods in computer vision. Specifically, the main contribution of this paper is identifying and interpreting the Transformer length extrapolation problem as a result of the limited expressiveness of the naive query and key dot product, and we successfully translate the length extrapolation issue into a well-understood feature map processing problem. The novel insight, which can be adapted to various attention-related models, reveals that the current Transformer architecture has the potential for further evolution. Extensive experiments demonstrate that treating attention as a feature map and applying convolution as a processing method significantly enhances Transformer performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan 等NeurIPS 2025 · 被引用 36 次
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodeSitong Wu, Haoru Tan, Yukang Chen, Shaofeng Zhang 等ICCV 2025 · 被引用 4 次
- SAS: Simulated Attention ScoreChuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang 等NeurIPS 2025 · 被引用 3 次
- Context-aware Biases for Length ExtrapolationAli Veisi, Hamidreza Amirzadeh, Amir MansourianEMNLP 2025 · 被引用 2 次
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMsXiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- Evolving Attention with Residual ConvolutionsYujing Wang, Yaming Yang, Jiangang Bai, Mingliang Zhang 等ICML 2021 · 被引用 43 次
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma 等ACL 2023 · 被引用 45 次
- Jump Self-attention: Capturing High-order Statistics in TransformersHaoyi Zhou, Siyang Xiao, Shanghang Zhang, Jieqi Peng 等NeurIPS 2022 · 被引用 4 次
- Linear Log-Normal Attention with Unbiased ConcentrationYury Nahshan, Joseph Kampeas, Emir HalevaICLR 2024 · 被引用 13 次
- Extrapolation by Association: Length Generalization Transfer In TransformersZiyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak 等NeurIPS 2025 · 被引用 13 次
