Location Attention for Extrapolation to Longer Sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia Bruni
摘要
Neural networks are surprisingly good at interpolating and perform remarkably well when the training set examples resemble those in the test set. However, they are often unable to extrapolate patterns beyond the seen data, even when the abstractions required for such patterns are simple. In this paper, we first review the notion of extrapolation, why it is important, and how one could hope to tackle it. We then focus on a specific type of extrapolation, which is especially useful for natural language processing: generalization to sequences longer than those seen during training. We hypothesize that models with a separate contentand location-based attention are more likely to extrapolate than those with common attention mechanisms. We empirically support our claim for recurrent seq2seq models with our proposed attention on variants of the Lookup Table task . This sheds light on some striking failures of neural models for sequences and on possible methods to approaching such issues.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic GeneralizationRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberICLR 2022 · 被引用 70 次
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 被引用 69 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
相关 Paper
- Monotonic Location Attention for Length GeneralizationJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 被引用 11 次
- How Do Neural Sequence Models Generalize? Local and Global Cues for Out-of-Distribution PredictionD. Anthony Bau, Jacob AndreasEMNLP 2021 · 被引用 2 次
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma 等ACL 2023 · 被引用 45 次
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention ModelsJiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li 等ICLR 2026 · 被引用 6 次
- Quantitative Bounds for Length Generalization in TransformersZachary Izzo, Eshaan Nichani, Jason D. LeeICLR 2026 · 被引用 8 次
