SSFT: Algorithm and Hardware Co-design for Structured Sparse Fine-Tuning of Large Language Models
Miao Yu, Trevor E. Carlson
Abstract
A significant number of users depend on Large Language Models (LLMs) for downstream tasks, but training LLMs from scratch remains prohibitively expensive. Sparse finetuning (SFT) has emerged as an effective strategy to reduce both the time and memory requirements of fine-tuning LLMs, achieving accuracy on par with fully fine-tuned models. Although SFT has the potential to achieve superior performance by minimizing computational requirements, SFT on GPUs often underperforms compared to dense algorithms like LoRA due to sparse data accesses that modern GPUs cannot efficiently handle. To address these issues, we propose Structured Sparse FineTuning (SSFT). It comprises a novel algorithm, SSFT-Alg, which introduces predictable sparsity patterns to reduce memory access overhead and enhance regularity in the SFT process. To support SSFT-Alg, an accelerator, SSFT-Hw, is proposed to optimize SSFT-Alg through an innovative sparsity-aware design, avoiding the overhead of sparsity operations on GPUs and optimizing latency and energy efficiency. Experiments with relevant models and benchmarks demonstrate that SSFT achieves comparable accuracy to state-of-the-art models on BERT, LLaMA 2 7B, and LLaMA 2 13B. Moreover, SSFT-Hw outperforms both GPUs and the state-of-the-art sparsity-aware transformer accelerators in throughput by and , respectively, while improving energy efficiency by and .
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- S2FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured SparsityXinyu Yang, Jixuan Leng, Geyang Guo, Jiawei Zhao et al.NeurIPS 2024 · 13 citations
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
- SMT: Fine-Tuning Large Language Models with Sparse MatricesHaoze He, Juncheng B. Li, Xuan Jiang, Heather MillerICLR 2025
- Efficient Transformer Inference with Statically Structured Sparse AttentionSteve Dai, Hasan Genc, Rangharajan Venkatesan, Brucek KhailanyDAC 2023 · 11 citations
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
