Over-parameterized Model Optimization with Polyak-Łojasiewicz Condition
Yixuan Chen, Yubin Shi, Mingzhi Dong, Xiaochen Yang, Dongsheng Li, Yujiang Wang, Robert P. Dick, Qin Lv, Yingying Zhao, Fan Yang, Ning Gu, Li Shang
Abstract
This work pursues the optimization of over-parameterized deep models for superior training efficiency and test performance. We first theoretically emphasize the importance of two properties of over-parameterized models, i.e., the convergence gap and the generalization gap. Subsequent analyses unveil that these two gaps can be upper-bounded by the ratio of the Lipschitz constant and the Polyak-Łojasiewicz (PL) constant, a crucial term abbreviated as the condition number. Such discoveries have led to a structured pruning method with a novel pruning criterion. That is, we devise a gating network that dynamically detects and masks out those poorly-behaved nodes of a deep model during the training session. To this end, this gating network is learned via minimizing the condition number of the target model, and this process can be implemented as an extra regularization loss term. Experimental studies demonstrate that the proposed method outperforms the baselines in terms of both training efficiency and test performance, exhibiting the potential of generalizing to a variety of deep network architectures and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59bcb359-f91d-4e18-996b-90cae8f74660Cited by top-tier papers2
- Understanding and Improving Training-free Loss-based Diffusion GuidanceYifei Shen, Xinyang Jiang, Yifan Yang, Yezhen Wang et al.NeurIPS 2024 · 36 citations
- SD-MoE: Spectral Decomposition for Effective Expert SpecializationRuijun Huang, Fang DONG(董方), Xin Zhang, Anrui Chen et al.ICML 2026 · 2 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- The Generalization-Stability Tradeoff In Neural Network PruningBrian R. Bartoldson, Ari S. Morcos, Adrian Barbu, Gordon ErlebacherNeurIPS 2020 · 97 citations
- Differentiable Sparsity via -Gating: Simple and Versatile Structured PenalizationChris Kolb, Laetitia Frost, Bernd Bischl, David RügamerNeurIPS 2025 · 4 citations
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 58 citations
- Pruning's Effect on Generalization Through the Lens of Training and RegularizationTian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle et al.NeurIPS 2022 · 40 citations
- Loss Landscape Characterization of Neural Networks without Over-ParametrizationRustem Islamov, Niccolò Ajroldi, Antonio Orvieto, Aurélien LucchiNeurIPS 2024 · 14 citations
