S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
Xinyu Yang, Jixuan Leng, Geyang Guo, Jiawei Zhao, Ryumei Nakada, Linjun Zhang, Huaxiu Yao, Beidi Chen
摘要
Current PEFT methods for LLMs can achieve high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability. Utilizing this key insight, we propose a family of Structured Sparse Fine-Tuning (S 2 FT) methods for LLMs, which concurrently achieve state-of-theart fine-tuning performance, training efficiency, and inference scalability. S 2 FT accomplishes this by "selecting sparsely and computing densely". Based on the coupled structures in LLMs, S 2 FT selects a few attention heads and channels in the MHA and FFN modules for each Transformer block, respectively. Next, it co-permutes the weight matrices on both sides of all coupled structures to connect the selected subsets in each layer into a dense submatrix. Finally, S 2 FT performs in-place gradient updates on all selected submatrices. Through theoretical analyses and empirical results, our method prevents forgetting while simplifying optimization, delivers SOTA performance on both commonsense and arithmetic reasoning with 4.6% and 1.3% average improvements compared to LoRA, and surpasses full FT by 11.5% when generalizing to various domains after instruction tuning. Using our partial back-propagation algorithm, S 2 FT saves training memory up to 3× and improves latency by 1.5-2.7× compared to full FT, while achieving an average 10% improvement over LoRA on both metrics. We further demonstrate that the weight updates in S 2 FT can be decoupled into adapters, enabling effective fusion, fast switch, and efficient parallelism when serving multiple fine-tuned models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax 等NeurIPS 2025 · 被引用 11 次
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded UpdatesAtsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos AletrasACL 2026 · 被引用 3 次
- MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing StructureJiale Kang, Qingyu YinICLR 2026 · 被引用 3 次
- DiaBlo: Diagonal Blocks Are Sufficient For FinetuningSelcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong 等ICLR 2026 · 被引用 2 次
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual SparsitySamir Khaki, Xiuyu Li, Junxian Guo, Ligeng Zhu 等ICML 2025
它引用的顶会 Paper27
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos 等ICLR 2024 · 被引用 433 次
相关 Paper
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and OptimizationXujia Wang, Yunjia Qi, Bin XuEMNLP 2025
- SSFT: Algorithm and Hardware Co-design for Structured Sparse Fine-Tuning of Large Language ModelsMiao Yu, Trevor E. CarlsonDAC 2025 · 被引用 1 次
- RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationMahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan AlistarhICML 2024 · 被引用 53 次
- SMT: Fine-Tuning Large Language Models with Sparse MatricesHaoze He, Juncheng B. Li, Xuan Jiang, Heather MillerICLR 2025
- Expanding Sparse Tuning for Low Memory UsageShufan Shen, Junshu Sun, Xiangyang Ji, Qingming Huang 等NeurIPS 2024 · 被引用 12 次
