Alignment-Enhanced Integration of Connectivity and Spectral Sparsity in Dynamic Sparse Training of LLM
Wenjing Wu, Yingtao Zhang, Jialin Zhao, Carlo Vittorio Cannistraci
摘要
With the rapid development of large language models (LLMs), identifying efficient strategies for training such large-scale systems has become increasingly critical. Although LLMs have achieved remarkable success across diverse applications, the necessity of maintaining full dense matrices during pre-training has been questioned, giving rise to parameter-efficient sparse pre-training methods which retains parameter-efficiency in both training and inference. These methods can be further divided into connectivity sparse training and spectral sparse training, with dynamic connectivity sparse training and low-rank factorization emerging as representative approaches for the two branches. However, a unified framework that effectively combines the strengths of both has yet to be established. In this work, we observe that the cancellation effect between the sparse and low-rank branches may limit the expressivity of the model, manifesting as output conflicts when the two components are combined. To address this issue, we first quantify the cancellation effect using the overlap cancellation ratio (OCR) and then propose a novel scheme that integrates dynamic sparse training with low-rank training, introducing a simple yet effective alignment loss to mitigate the disagreement between the two branches and promote better collaboration. We validate this scheme by combining a representative dynamic sparse training method, CHTs, with low-rank training, resulting in a new parameter-efficient training approach termed CHTsL. The method is evaluated on LLaMA60M and LLaMA130M using the OpenWebText and C4 datasets, where only 10%, 20%, and 30% of the parameters are preserved compared to dense training. Experimental results demonstrate that our proposed scheme effectively alleviates the cancellation effect, especially in the Q and K matrices of the attention layers, and improves training stability and performance compared to the naive combination of sparse and low-rank components. Additionally, the new scheme enables CHTsL to consistently outperform other parameter-efficient sparse training methods under the same parameter budget, achieving performance closest to that of dense training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 被引用 308 次
相关 Paper
- SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language ModelsXudong Lu, Aojun Zhou, Yuhui Xu, Renrui Zhang 等ICML 2024 · 被引用 14 次
- Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical ScalingQiao Xiao, Boqian Wu, Patrik Okanovic, Tomasz Sternal 等ICML 2026
- SLTrain: a sparse plus low rank approach for parameter and memory efficient pretrainingAndi Han, Jiaxiang Li, Wei Huang, Mingyi Hong 等NeurIPS 2024 · 被引用 54 次
- The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling LawsTian Jin, Ahmed Imtiaz Humayun, Utku Evci, Suvinay Subramanian 等ICLR 2025
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao 等NeurIPS 2025 · 被引用 17 次
