Lune

NeurIPS2025Top-tier venue

Multi-Objective One-Shot Pruning for Large Language Models

Weiyu Chen, Hansi Yang, Yunhao Gou, Han Shi, Enliang Hu, Zhenguo Li, James Kwok

2025Year

Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks but require substantial computational resources, limiting their deployment in resource-constrained environments. While one-shot pruning methods can reduce model size without expensive retraining, they typically optimize for single objectives, ignoring LLMs' multi-faceted applications. We introduce Multi-Objective One-Shot Pruning (MOSP), which formulates LLM pruning as a multi-objective optimization problem. MOSP efficiently generates a Pareto set of pruned models representing different capability trade-offs, allowing users to select solutions aligned with their preferences. The proposed approach identifies share core support while enabling specialized support. Experiments across various LLMs and sparsity levels demonstrate MOSP's superior performance in navigating multi-objective trade-offs compared to baseline methods.

consider calibration data X from a single, general-purpose dataset. This overlooks the fact that LLMs are often evaluated across multiple criteria. Different users may prioritize these objectives differently. Current one-shot pruning techniques generally do not address this need for customization, lacking mechanisms to generate models tailored to specific user preferences.

Another line of work considers the allocation of sparsity across different layers [41,45,23,39]. Such layer-wise sparsity distribution strategies can often be combined with most of the aforementioned pruning algorithms to further improve performance. These approaches are orthogonal to the proposed methods and the two can be combined in a straightforward manner.

Multi-objective optimization (MOO) [31] optimizes m objective functions simultaneously. Without loss of generality, we consider the minimization problem: min θ∈Θ f (θ) = min θ∈Θ f 1 (θ), . . . , f m (θ) , where Θ is the feasible decision space. A solution a dominates b, denoted a ≺ b, if ∀i ∈ 1, . . . , m : f i (a) ≤ f i (b) and ∃j ∈ 1, . . . , m : f j (a) < f j (b). A feasible solution is Pareto-optimal when it is not dominated by any other feasible solution. The set of all Pareto-optimal decision vectors is called the Pareto set. The corresponding set of objective vectors, F * = f (θ) | θ is Pareto-optimal, is the Pareto front. Gradient-based MOO methods have been widely adopted in deep learning [8]. They can be classified into three main categories: (i) Learning a single solution, with examples including MGDA [32, 14, 12], CAGrad [25], and Nash-MTL [33]; (ii) Learning a finite Pareto set, with examples including PMTL[24], EPO [29], MOO-SVGD [26], and GMOOAR [6]; and (3) Learning an infinite set of solutions, with examples including PHN [34], PaMaL [13], and LORPMAN [7].

All the aforementioned algorithms utilize gradient descent for optimization. However, the direct application of gradient descent is empirically ineffective in obtaining satisfactory solutions in the unstructured LLM pruning scenario, particularly when dealing with high sparsity ratios [30]. Consequently, these algorithms are not directly amenable to modification for one-shot LLM pruning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6aaa95cf-b861-4575-95e9-cfe063b1f2ce

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines