OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
Xiang Meng, Shibal Ibrahim, Kayhan Behdin, Hussein Hazimeh, Natalia Ponomareva, Rahul Mazumder
摘要
Structured pruning is a promising approach for reducing the inference costs of large vision and language models. By removing carefully chosen structures, e.g., neurons or attention heads, the improvements from this approach can be realized on standard deep learning hardware. In this work, we focus on structured pruning in the one-shot (post-training) setting, which does not require model retraining after pruning. We propose a novel combinatorial optimization framework for this problem, based on a layer-wise reconstruction objective and a careful reformulation that allows for scalable optimization. Moreover, we design a new local combinatorial optimization algorithm, which exploits low-rank updates for efficient local search. Our framework is time and memory-efficient and considerably improves upon state-of-the-art one-shot methods on vision models (e.g., ResNet50, MobileNet) and language models (e.g., OPT-1.3B -- OPT-30B). For language models, e.g., OPT-2.7B, OSSCAR can lead to lower test perplexity on WikiText with inference time speedup in comparison to the state-of-the-art ZipLM approach. Our framework is also -- faster. Notably, our work considers models with tens of billions of parameters, which is up to larger than what has been previously considered in the structured pruning literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language ModelsHaiquan Lu, Yefan Zhou, Shiwei Liu, Zhangyang Wang 等NeurIPS 2024 · 被引用 49 次
- ALPS: Improved Optimization for Highly Sparse One-Shot Pruning for Large Language ModelsXiang Meng, Kayhan Behdin, Haoyue Wang, Rahul MazumderNeurIPS 2024 · 被引用 19 次
- Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution OptimizationGuanchen Li, Yixing Xu, Zeping Li, Ji Liu 等NeurIPS 2025 · 被引用 7 次
- Free Lunch in the Forest: Functionally-Identical Pruning of Boosted Tree EnsemblesYoussouf Emine, Alexandre Forel, Idriss Malek, Thibaut VidalAAAI 2025 · 被引用 3 次
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMsMehdi Makni, Xiang Meng, Rahul MazumderNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper8
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian 等NeurIPS 2023 · 被引用 495 次
- A Fast Post-Training Pruning Framework for TransformersWoosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun 等NeurIPS 2022 · 被引用 247 次
- Group Fisher Pruning for Practical Network CompressionLiyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou 等ICML 2021 · 被引用 204 次
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan 等NeurIPS 2021 · 被引用 198 次
相关 Paper
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- Structured Optimal Brain Pruning for Large Language ModelsJiateng Wei, Quan Lu, Ning Jiang, Siqi Li 等EMNLP 2024 · 被引用 2 次
- The LLM SurgeonTycho F. A. van der Ouderaa, Markus Nagel, Mart van Baalen, Tijmen BlankevoortICLR 2024 · 被引用 29 次
- A Robust Optimization Guided Pruning Framework for Vision and Large Language ModelsGabriel Afriat, Hussein Hazimeh, Dimitris Paparas, Rahul MazumderICML 2026
- Preserving Deep Representations in One-Shot Pruning: A Hessian-Free Second-Order Optimization FrameworkRyan Lucas, Rahul MazumderICLR 2025
