Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
Hanseul Cho, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun
Abstract
Transformers often struggle with length generalization, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commonly used to study length generalization, certain tasks are considered notoriously difficult, e.g., multi-operand addition (requiring generalization over both the number of operands and their lengths) and multiplication (requiring generalization over both operand lengths). In this work, we achieve approximately 2-3x length generalization on both tasks, which is the first such achievement in arithmetic Transformers. We design task-specific scratchpads enabling the model to focus on a fixed number of tokens per each next-token prediction step, and apply multi-level versions of Coupling (Cho et al., 2024; McLeish et al., 2024) to let Transformers know the right position to attend to. On the theory side, we prove that a 1-layer Transformer using our method can solve multi-operand addition, up to operand length and operand count that are exponential in embedding dimension.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c45c1116-6007-4bfa-b64a-6855fa5297f6Cited by top-tier papers11
- Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task StructureHanseul Cho, Jaeyoung Cha, Pranjal Awasthi, Srinadh Bhojanapalli et al.NeurIPS 2024 · 28 citations
- Quantitative Bounds for Length Generalization in TransformersZachary Izzo, Eshaan Nichani, Jason D. LeeICLR 2026 · 8 citations
- Beyond Single-Task: Robust Multi-Task Length Generalization for LLMsYi Hu, Shijia Kang, Haotong Yang, Haotian Xu et al.NeurIPS 2025 · 6 citations
- Characterizing Pattern Matching and Its Limits on Compositional Task StructuresHoyeon Chang, Jinho Park, Hanseul Cho, Sohee Yang et al.ICLR 2026 · 5 citations
- Circuit Stability Characterizes Language Model GeneralizationAlan SunACL 2025 · 4 citations
Related papers
- Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning TasksXingcheng Xu, Zibo Zhao, Haipeng Zhang, Yanqing YangACL 2025 · 2 citations
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade et al.ICML 2025
- Transformers Can Do Arithmetic with the Right EmbeddingsSean McLeish, Arpit Bansal, Alex Stein, Neel Jain et al.NeurIPS 2024 · 94 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- Looped Transformers for Length GeneralizationYing Fan, Yilun Du, Kannan Ramchandran, Kangwook LeeICLR 2025
