S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
Xinyu Yang, Jixuan Leng, Geyang Guo, Jiawei Zhao, Ryumei Nakada, Linjun Zhang, Huaxiu Yao, Beidi Chen
Abstract
Current PEFT methods for LLMs can achieve high quality, efficient training, or scalable serving, but not all three simultaneously. To address this limitation, we investigate sparse fine-tuning and observe a remarkable improvement in generalization ability. Utilizing this key insight, we propose a family of Structured Sparse Fine-Tuning (S 2 FT) methods for LLMs, which concurrently achieve state-of-theart fine-tuning performance, training efficiency, and inference scalability. S 2 FT accomplishes this by "selecting sparsely and computing densely". Based on the coupled structures in LLMs, S 2 FT selects a few attention heads and channels in the MHA and FFN modules for each Transformer block, respectively. Next, it co-permutes the weight matrices on both sides of all coupled structures to connect the selected subsets in each layer into a dense submatrix. Finally, S 2 FT performs in-place gradient updates on all selected submatrices. Through theoretical analyses and empirical results, our method prevents forgetting while simplifying optimization, delivers SOTA performance on both commonsense and arithmetic reasoning with 4.6% and 1.3% average improvements compared to LoRA, and surpasses full FT by 11.5% when generalizing to various domains after instruction tuning. Using our partial back-propagation algorithm, S 2 FT saves training memory up to 3× and improves latency by 1.5-2.7× compared to full FT, while achieving an average 10% improvement over LoRA on both metrics. We further demonstrate that the weight updates in S 2 FT can be decoupled into adapters, enabling effective fusion, fast switch, and efficient parallelism when serving multiple fine-tuned models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afcdac1e-67f7-41fe-938d-a0841c39bb4bCited by top-tier papers9
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax et al.NeurIPS 2025 · 11 citations
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded UpdatesAtsuki Yamaguchi, Terufumi Morishita, Aline Villavicencio, Nikolaos AletrasACL 2026 · 3 citations
- MiSS: Revisiting the Trade-off in LoRA with an Efficient Shard-Sharing StructureJiale Kang, Qingyu YinICLR 2026 · 3 citations
- DiaBlo: Diagonal Blocks Are Sufficient For FinetuningSelcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong et al.ICLR 2026 · 2 citations
- SparseLoRA: Accelerating LLM Fine-Tuning with Contextual SparsitySamir Khaki, Xiuyu Li, Junxian Guo, Ligeng Zhu et al.ICML 2025
Builds on27
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 · 433 citations
Related papers
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and OptimizationXujia Wang, Yunjia Qi, Bin XuEMNLP 2025
- SSFT: Algorithm and Hardware Co-design for Structured Sparse Fine-Tuning of Large Language ModelsMiao Yu, Trevor E. CarlsonDAC 2025 · 1 citation
- RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationMahdi Nikdan, Soroush Tabesh, Elvir Crncevic, Dan AlistarhICML 2024 · 53 citations
- SMT: Fine-Tuning Large Language Models with Sparse MatricesHaoze He, Juncheng B. Li, Xuan Jiang, Heather MillerICLR 2025
- Expanding Sparse Tuning for Low Memory UsageShufan Shen, Junshu Sun, Xiangyang Ji, Qingming Huang et al.NeurIPS 2024 · 12 citations
