Multi-Objective One-Shot Pruning for Large Language Models
Weiyu Chen, Hansi Yang, Yunhao Gou, Han Shi, Enliang Hu, Zhenguo Li, James Kwok
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks but require substantial computational resources, limiting their deployment in resource-constrained environments. While one-shot pruning methods can reduce model size without expensive retraining, they typically optimize for single objectives, ignoring LLMs' multi-faceted applications. We introduce Multi-Objective One-Shot Pruning (MOSP), which formulates LLM pruning as a multi-objective optimization problem. MOSP efficiently generates a Pareto set of pruned models representing different capability trade-offs, allowing users to select solutions aligned with their preferences. The proposed approach identifies share core support while enabling specialized support. Experiments across various LLMs and sparsity levels demonstrate MOSP's superior performance in navigating multi-objective trade-offs compared to baseline methods.
consider calibration data X from a single, general-purpose dataset. This overlooks the fact that LLMs are often evaluated across multiple criteria. Different users may prioritize these objectives differently. Current one-shot pruning techniques generally do not address this need for customization, lacking mechanisms to generate models tailored to specific user preferences.
Another line of work considers the allocation of sparsity across different layers [41,45,23,39]. Such layer-wise sparsity distribution strategies can often be combined with most of the aforementioned pruning algorithms to further improve performance. These approaches are orthogonal to the proposed methods and the two can be combined in a straightforward manner.
Multi-objective optimization (MOO) [31] optimizes m objective functions simultaneously. Without loss of generality, we consider the minimization problem: min θ∈Θ f (θ) = min θ∈Θ f 1 (θ), . . . , f m (θ) , where Θ is the feasible decision space. A solution a dominates b, denoted a ≺ b, if ∀i ∈ 1, . . . , m : f i (a) ≤ f i (b) and ∃j ∈ 1, . . . , m : f j (a) < f j (b). A feasible solution is Pareto-optimal when it is not dominated by any other feasible solution. The set of all Pareto-optimal decision vectors is called the Pareto set. The corresponding set of objective vectors, F * = f (θ) | θ is Pareto-optimal, is the Pareto front. Gradient-based MOO methods have been widely adopted in deep learning [8]. They can be classified into three main categories: (i) Learning a single solution, with examples including MGDA [32, 14, 12], CAGrad [25], and Nash-MTL [33]; (ii) Learning a finite Pareto set, with examples including PMTL[24], EPO [29], MOO-SVGD [26], and GMOOAR [6]; and (3) Learning an infinite set of solutions, with examples including PHN [34], PaMaL [13], and LORPMAN [7].
All the aforementioned algorithms utilize gradient descent for optimization. However, the direct application of gradient descent is empirically ineffective in obtaining satisfactory solutions in the unstructured LLM pruning scenario, particularly when dealing with high sparsity ratios [30]. Consequently, these algorithms are not directly amenable to modification for one-shot LLM pruning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aaa95cf-b861-4575-95e9-cfe063b1f2ceBuilds on22
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
Related papers
- ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language ModelsXiang Meng, Kayhan Behdin, Haoyue Wang, Rahul MazumderNeurIPS 2024 · 19 citations
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- A Robust Optimization Guided Pruning Framework for Vision and Large Language ModelsGabriel Afriat, Hussein Hazimeh, Dimitris Paparas, Rahul MazumderICML 2026
- Lua-LLM: Learning Unstructured-Sparsity Allocation for Large Language ModelsMingge Lu, Jingwei Sun, Junqing Lin, Zechun Zhou et al.NeurIPS 2025 · 1 citation
- M-Wanda: Improving One-Shot Pruning for Multilingual LLMsRochelle Choenni, Ivan TitovEMNLP 2025 · 1 citation
