MVA: Linear Attention with High-order Query-Keys Integration and Multi-level Vocabulary Decomposition
Ning Wang, Zekun Li, Tongxin Bai, Man Yao, Zhen Qin, Guoqi Li
Abstract
Linear attention offers the advantages of linear inference time and fixed memory usage compared to Softmax attention. However, training large-scale language models with linear attention from scratch remains prohibitively expensive and exhibits significant performance gaps compared to Softmax-based models. To address these challenges, we focus on transforming pre-trained Softmax-based language models into linear attention models. We unify mainstream linear attention methods using a high-order QK integration theory and a multi-level vocabulary decomposition. Specifically, the QK integration theory explains the efficacy of combining linear and sparse attention from the perspective of information collection across different frequency bands. The multi-level vocabulary decomposition exponentially expands memory capacity by recursively exploiting compression loss from compressed states. Through detailed error analysis, we demonstrate superior approximation of Softmax attention achieved by our approach. To further improve performance and reduce training costs, we adopt a soft integration strategy with attention scores, effectively combining a sliding window mechanism. With less than 100M tokens, our method fine-tunes models to achieve linear complexity while retaining 99% of their original performance. Compared to state-of-the-art linear attention model and method, our approach improves MMLU scores by 1.2 percentage points with minimal fine-tuning. Furthermore, even without the sliding window mechanism, our method achieves state-of-the-art performance on all test sets with 10B tokens.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large ModelsJunkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian et al.AAAI 2026 · 12 citations
- SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free AccelerationZekun Li, Ning Wang, Tongxin Bai, Changwang Mei et al.CVPR 2026 · 4 citations
Builds on19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Diagonal State Spaces are as Effective as Structured State SpacesAnkit Gupta, Albert Gu, Jonathan BerantNeurIPS 2022 · 546 citations
Related papers
- Luna: Linear Unified Nested AttentionXuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou et al.NeurIPS 2021 · 145 citations
- Native Hybrid Attention for Efficient Sequence ModelingJusen Du, Jiaxi Hu, Zhang Tao, Weigao Sun et al.ACL 2026 · 8 citations
- Lizard: An Efficient Linearization Framework for Large Language ModelsChien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy et al.ACL 2026 · 8 citations
- Dynamic Linear AttentionXin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu et al.ICML 2026 · 1 citation
- Rectifying Magnitude Neglect in Linear AttentionQihang Fan, Huaibo Huang, Yuang Ai, Ran HeICCV 2025 · 14 citations
