Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
Hanseul Cho, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun
摘要
Transformers often struggle with length generalization, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commonly used to study length generalization, certain tasks are considered notoriously difficult, e.g., multi-operand addition (requiring generalization over both the number of operands and their lengths) and multiplication (requiring generalization over both operand lengths). In this work, we achieve approximately 2-3x length generalization on both tasks, which is the first such achievement in arithmetic Transformers. We design task-specific scratchpads enabling the model to focus on a fixed number of tokens per each next-token prediction step, and apply multi-level versions of Coupling (Cho et al., 2024; McLeish et al., 2024) to let Transformers know the right position to attend to. On the theory side, we prove that a 1-layer Transformer using our method can solve multi-operand addition, up to operand length and operand count that are exponential in embedding dimension.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task StructureHanseul Cho, Jaeyoung Cha, Pranjal Awasthi, Srinadh Bhojanapalli 等NeurIPS 2024 · 被引用 28 次
- Quantitative Bounds for Length Generalization in TransformersZachary Izzo, Eshaan Nichani, Jason D. LeeICLR 2026 · 被引用 8 次
- Beyond Single-Task: Robust Multi-Task Length Generalization for LLMsYi Hu, Shijia Kang, Haotong Yang, Haotian Xu 等NeurIPS 2025 · 被引用 6 次
- Characterizing Pattern Matching and Its Limits on Compositional Task StructuresHoyeon Chang, Jinho Park, Hanseul Cho, Sohee Yang 等ICLR 2026 · 被引用 5 次
- Circuit Stability Characterizes Language Model GeneralizationAlan SunACL 2025 · 被引用 4 次
相关 Paper
- Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning TasksXingcheng Xu, Zibo Zhao, Haipeng Zhang, Yanqing YangACL 2025 · 被引用 2 次
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade 等ICML 2025
- Transformers Can Do Arithmetic with the Right EmbeddingsSean McLeish, Arpit Bansal, Alex Stein, Neel Jain 等NeurIPS 2024 · 被引用 94 次
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Looped Transformers for Length GeneralizationYing Fan, Yilun Du, Kannan Ramchandran, Kangwook LeeICLR 2025
