SAFE: Finding Sparse and Flat Minima to Improve Pruning
Dongyeop Lee, Kwanhee Lee, Jinseok Chung, Namhoon Lee
Abstract
Sparsifying neural networks often suffers from seemingly inevitable performance degradation, and it remains challenging to restore the original performance despite much recent progress. Motivated by recent studies in robust optimization, we aim to tackle this problem by finding subnetworks that are both sparse and flat at the same time. Specifically, we formulate pruning as a sparsity-constrained optimization problem where flatness is encouraged as an objective. We solve it explicitly via an augmented Lagrange dual approach and extend it further by proposing a generalized projection operation, resulting in novel pruning methods called SAFE and its extension, SAFE + . Extensive evaluations on standard image classification and language modeling tasks reveal that SAFE consistently yields sparse networks with improved generalization performance, which compares competitively to well-established baselines. In addition, SAFE demonstrates resilience to noisy data, making it well-suited for real-world conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c1dc467-7281-44a3-b8f6-de7cd88bc50dCited by top-tier papers2
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMsGunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim et al.ICLR 2026 · 8 citations
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMMKwanhee Lee, Hyeondo Jang, Dongyeop Lee, Dan Alistarh et al.ICLR 2026 · 5 citations
Builds on29
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
Related papers
- CSTAR: Towards Compact and Structured Deep Neural Networks with Adversarial RobustnessHuy Phan, Miao Yin, Yang Sui, Bo Yuan et al.AAAI 2023 · 10 citations
- HYDRA: Pruning Adversarially Robust Neural NetworksVikash Sehwag, Shiqi Wang, Prateek Mittal, Suman JanaNeurIPS 2020 · 242 citations
- Adaptive Sharpness-Aware Pruning for Robust Sparse NetworksAnna Bair, Hongxu Yin, Maying Shen, Pavlo Molchanov et al.ICLR 2024 · 19 citations
- A Robust Optimization Guided Pruning Framework for Vision and Large Language ModelsGabriel Afriat, Hussein Hazimeh, Dimitris Paparas, Rahul MazumderICML 2026
- Controlled Sparsity via Constrained Optimization or: How I Learned to Stop Tuning Penalties and Love ConstraintsJose Gallego-Posada, Juan Ramirez, Akram Erraqabi, Yoshua Bengio et al.NeurIPS 2022 · 32 citations
