Scaling Laws for Linear Complexity Language Models
Xuyang Shen, Dong Li, Ruitao Leng, Zhen Qin, Weigao Sun, Yiran Zhong
摘要
The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability. Specifically, we examine the scaling behaviors of three efficient linear architectures. These include TNL (Qin et al., 2024c), a linear attention model with data-independent decay; HGRN2 (Qin et al., 2024e), a linear RNN with data-dependent decay; and cosFormer2 (Qin et al., 2022b(Qin et al., , 2024a)), a linear attention model without decay. We also include LLaMA as a baseline architecture for comparison with softmax attention. These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks. These tasks include validation loss, commonsense reasoning, and information retrieval and generation. The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- MoM: Linear Sequence Modeling with Mixture-of-MemoriesJusen Du, Weigao Sun, Disen Lan, Jiaxi Hu 等ICLR 2026 · 被引用 43 次
- MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapYuhong Chou, Man Yao, Kexin Wang, Yuqi Pan 等NeurIPS 2024 · 被引用 22 次
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
- xLSTM Scaling Laws: Competitive Performance with Linear Time-ComplexityMaximilian Beck, Kajetan Schweighofer, Sebastian Böck, Sebastian Lehner 等ICLR 2026 · 被引用 3 次
- Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersWeigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu MoAAAI 2025 · 被引用 3 次
它引用的顶会 Paper4
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li 等ICLR 2022 · 被引用 303 次
- Hierarchically Gated Recurrent Neural Network for Sequence ModelingZhen Qin, Songlin Yang, Yiran ZhongNeurIPS 2023 · 被引用 152 次
- Linear Complexity Randomized Self-attention MechanismLin Zheng, Chong Wang, Lingpeng KongICML 2022 · 被引用 39 次
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 被引用 2 次
相关 Paper
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda 等ICML 2024 · 被引用 390 次
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk ParkICLR 2026 · 被引用 3 次
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang 等ACL 2026 · 被引用 17 次
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
