Implicit Language Models are RNNs: Balancing Parallelization and Expressivity
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani, Jannes Gladrow
摘要
State-space models (SSMs) and transformers dominate the language modeling landscape. However, they are constrained to a lower computational complexity than classical recurrent neural networks (RNNs), limiting their expressivity. In contrast, RNNs lack parallelization during training, raising fundamental questions about the trade off between parallelization and expressivity. We propose implicit SSMs, which iterate a transformation until convergence to a fixed point. Theoretically, we show that implicit SSMs implement the non-linear state-transitions of RNNs. Empirically, we find that only approximate fixed-point convergence suffices, enabling the design of a scalable training curriculum that largely retains parallelization, with full convergence required only for a small subset of tokens. Our approach demonstrates superior state-tracking capabilities on regular languages, surpassing transformers and SSMs. We further scale implicit SSMs to natural language reasoning tasks and pretraining of large-scale language models up to 1.3B parameters on 207B tokensrepresenting, to our knowledge, the largest implicit model trained to date. Notably, our implicit models outperform their explicit counterparts on standard benchmarks. Our code is publicly available at github.com/microsoft/ implicit_languagemodels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder ProductsJulien Siems, Timur Carstensen, Arber Zela, Frank Hutter 等NeurIPS 2025 · 被引用 75 次
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online OptimizationAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniICLR 2026 · 被引用 63 次
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri 等ICML 2026 · 被引用 57 次
- MesaNet: Sequence Modeling by Locally Optimal Test-Time TrainingJohannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari 等ICLR 2026 · 被引用 45 次
它引用的顶会 Paper29
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen 等NeurIPS 2024 · 被引用 412 次
相关 Paper
- ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language ModelsFederico Danieli, Pau Rodríguez, Miguel Sarabia, Xavier Suau 等ICLR 2026 · 被引用 18 次
- The Expressive Capacity of State Space Models: A Formal Language PerspectiveYash Raj Sarrof, Yana Veitsman, Michael HahnNeurIPS 2024 · 被引用 53 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
- Language models can learn implicit multi-hop reasoning, but only if they have lots of training dataYuekun Yao, Yupei Du, Dawei Zhu, Michael Hahn 等EMNLP 2025
- Fixed-Point RNNs: Interpolating from Diagonal to DenseSajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, Antonio OrvietoNeurIPS 2025 · 被引用 5 次
