MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
Yulong Huang, Xiang Liu, Hongxiang Huang, Xiaopeng LIN, Zunchang LIU, Xiaowen Chu, Zeke Xie, Bojun Cheng
摘要
Linear Attention (LA) offers a promising paradigm for scaling large language models (LLMs) to long sequences by avoiding the quadratic complexity of self-attention. Recent LA models such as Mamba2 and GDN interpret linear recurrences as closed-form online stochastic gradient descent (SGD), but naive SGD updates suffer from rapid information decay and suboptimal convergence in optimization. While momentum-based optimizers provide a natural remedy, they pose challenges in simultaneously achieving training efficiency and effectiveness. To address this, we develop a chunkwise parallel algorithm for LA with a stepwise momentum rule by geometrically reordering the update coefficients. Further, from a dynamical systems perspective, we analyze the momentum-based recurrence as a second-order system that introduces complex conjugate eigenvalues. This analysis guides the design of stable gating constraints. The resulting model, Momentum DeltaNet (MDN), leverages Triton kernels to achieve comparable training throughput with competitive linear models such as Mamba2 and KDA. Extensive experiments on the 400M and 1.3B parameter models demonstrate consistent performance improvements over strong baselines, including Transformers, Mamba2 and GDN, across diverse downstream evaluation benchmarks. Code: https://github.com/HuuYuLong/MomentumDeltaNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache CompressionXiang Liu, Zhenheng Tang, Hong Chen, Peijie Dong 等ICML 2026 · 被引用 16 次
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM InferenceXiang Liu, Xuming Hu, Xiaowen Chu, Eunsol ChoiICLR 2026 · 被引用 15 次
它引用的顶会 Paper32
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear RecurrencesNeehal Tumma, Noel Loo, Daniela RusICML 2026 · 被引用 3 次
- Gated Delta Networks: Improving Mamba2 with Delta RuleSonglin Yang, Jan Kautz, Ali HatamizadehICLR 2025
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen 等NeurIPS 2024 · 被引用 412 次
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing 等ICLR 2026 · 被引用 41 次
- Gated KalmaNet: A Fading Memory Layer through Test-time Ridge RegressionLiangzu Peng, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez 等CVPR 2026 · 被引用 10 次
