Improving Bilinear RNN with Closed-loop Control
Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang, Weigao Sun
Abstract
Recent efficient sequence modeling methods such as Gated DeltaNet, TTT, and RWKV-7 have achieved performance improvements by supervising the recurrent memory management through Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, structurally resembling bilinear systems. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then, based on closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on large-scale corpus. Comba demonstrates superior performance and computation efficiency in both language and vision modeling. https://github.com/fla-org/flash-linear-attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8c7b1fd-57e8-40d3-9fe5-39a226b00f04Cited by top-tier papers4
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- Memory Caching: RNNs with Growing MemoryAli Behrouz, Zeman Li, Yuan Deng, Peilin Zhong et al.ICML 2026 · 10 citations
- MDN: Parallelizing Stepwise Momentum for Delta Linear AttentionYulong Huang, Xiang Liu, Hongxiang Huang, Xiaopeng LIN et al.ICML 2026 · 1 citation
- Push, Pop, Parallelize: Stack-Augmented Linear Attention via the Delta RuleAnh T Nguyen, Saleh Momeni, Ashutosh Chaubey, Changnan Xiao et al.ICML 2026
Builds on48
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
Related papers
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen et al.NeurIPS 2024 · 412 citations
- DeltaProduct: Improving State-Tracking in Linear RNNs via Householder ProductsJulien Siems, Timur Carstensen, Arber Zela, Frank Hutter et al.NeurIPS 2025 · 75 citations
- Gated KalmaNet: A Fading Memory Layer through Test-time Ridge RegressionLiangzu Peng, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez et al.CVPR 2026 · 10 citations
- Q-Delta: Beyond Key–Value Associative State EvolutionSumin Park, Seojin Kim, Noseong ParkICML 2026
- Gated Delta Networks: Improving Mamba2 with Delta RuleSonglin Yang, Jan Kautz, Ali HatamizadehICLR 2025
