Location Attention for Extrapolation to Longer Sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, Elia Bruni
Abstract
Neural networks are surprisingly good at interpolating and perform remarkably well when the training set examples resemble those in the test set. However, they are often unable to extrapolate patterns beyond the seen data, even when the abstractions required for such patterns are simple. In this paper, we first review the notion of extrapolation, why it is important, and how one could hope to tackle it. We then focus on a specific type of extrapolation, which is especially useful for natural language processing: generalization to sequences longer than those seen during training. We hypothesize that models with a separate contentand location-based attention are more likely to extrapolate than those with common attention mechanisms. We empirically support our claim for recurrent seq2seq models with our proposed attention on variants of the Lookup Table task . This sheds light on some striking failures of neural models for sequences and on possible methods to approaching such issues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d18b3197-b27a-44a3-97c6-9be92bb62618Cited by top-tier papers16
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic GeneralizationRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberICLR 2022 · 70 citations
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 69 citations
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie et al.ICLR 2024 · 66 citations
Related papers
- Monotonic Location Attention for Length GeneralizationJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 11 citations
- How Do Neural Sequence Models Generalize? Local and Global Cues for Out-of-Distribution PredictionD. Anthony Bau, Jacob AndreasEMNLP 2021 · 2 citations
- A Length-Extrapolatable TransformerYutao Sun, Li Dong, Barun Patra, Shuming Ma et al.ACL 2023 · 45 citations
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention ModelsJiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li et al.ICLR 2026 · 6 citations
- Quantitative Bounds for Length Generalization in TransformersZachary Izzo, Eshaan Nichani, Jason D. LeeICLR 2026 · 8 citations
