Scaling Laws for Linear Complexity Language Models
Xuyang Shen, Dong Li, Ruitao Leng, Zhen Qin, Weigao Sun, Yiran Zhong
Abstract
The interest in linear complexity models for large language models is on the rise, although their scaling capacity remains uncertain. In this study, we present the scaling laws for linear complexity language models to establish a foundation for their scalability. Specifically, we examine the scaling behaviors of three efficient linear architectures. These include TNL (Qin et al., 2024c), a linear attention model with data-independent decay; HGRN2 (Qin et al., 2024e), a linear RNN with data-dependent decay; and cosFormer2 (Qin et al., 2022b(Qin et al., , 2024a)), a linear attention model without decay. We also include LLaMA as a baseline architecture for comparison with softmax attention. These models were trained with six variants, ranging from 70M to 7B parameters on a 300B-token corpus, and evaluated with a total of 1,376 intermediate checkpoints on various downstream tasks. These tasks include validation loss, commonsense reasoning, and information retrieval and generation. The study reveals that existing linear complexity language models exhibit similar scaling capabilities as conventional transformer-based models while also demonstrating superior linguistic proficiency and knowledge retention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c89b115-4e80-41da-95e4-014523dc1905Cited by top-tier papers8
- MoM: Linear Sequence Modeling with Mixture-of-MemoriesJusen Du, Weigao Sun, Disen Lan, Jiaxi Hu et al.ICLR 2026 · 43 citations
- MetaLA: Unified Optimal Linear Approximation to Softmax Attention MapYuhong Chou, Man Yao, Kexin Wang, Yuqi Pan et al.NeurIPS 2024 · 22 citations
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro et al.ICLR 2026 · 7 citations
- xLSTM Scaling Laws: Competitive Performance with Linear Time-ComplexityMaximilian Beck, Kajetan Schweighofer, Sebastian Böck, Sebastian Lehner et al.ICLR 2026 · 3 citations
- Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersWeigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu MoAAAI 2025 · 3 citations
Builds on4
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li et al.ICLR 2022 · 303 citations
- Hierarchically Gated Recurrent Neural Network for Sequence ModelingZhen Qin, Songlin Yang, Yiran ZhongNeurIPS 2023 · 152 citations
- Linear Complexity Randomized Self-attention MechanismLin Zheng, Chong Wang, Lingpeng KongICML 2022 · 39 citations
- Efficient Attention via Control VariatesLin Zheng, Jianbo Yuan, Chong Wang, Lingpeng KongICLR 2023 · 2 citations
Related papers
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- Deriving Neural Scaling Laws from the Statistics of Natural LanguageFrancesco Cagnetta, Allan Raventos, Surya Ganguli, Matthieu WyartICML 2026
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk ParkICLR 2026 · 3 citations
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang et al.ACL 2026 · 17 citations
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge et al.ICML 2025
