Improving Generalization of Deep Neural Networks by Optimum Shifting
Yuyan Zhou, Ye Li, Lei Feng, Sheng-Jun Huang
Abstract
Recent studies showed that the generalization of neural networks is correlated with the sharpness of the loss landscape, and flat minima suggests a better generalization ability than sharp minima. In this paper, we propose a novel method called optimum shifting, which changes the parameters of a neural network from a sharp minimum to a flatter one while maintaining the same training loss value. Our method is based on the observation that when the input and output of a neural network are fixed, the matrix multiplications within the network can be treated as systems of under-determined linear equations, enabling adjustment of parameters in the solution space, which can be simply accomplished by solving a constrained optimization problem. Furthermore, we introduce a practical stochastic optimum shifting technique utilizing the Neural Collapse theory to reduce computational costs and provide more degrees of freedom for optimum shifting. Extensive experiments (including classification and detection) with various deep neural network architectures on benchmark datasets demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Penalizing Gradient Norm for Efficiently Improving Generalization in Deep LearningYang Zhao, Hao Zhang, Xiuyuan HuICML 2022 · 165 citations
- On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained FeaturesJinxin Zhou, Xiao Li, Tianyu Ding, Chong You et al.ICML 2022 · 122 citations
Related papers
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 53 citations
- Deep Neural Collapse Is Provably Optimal for the Deep Unconstrained Features ModelPeter Súkeník, Marco Mondelli, Christoph H. LampertNeurIPS 2023 · 51 citations
- GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved GeneralizationZhiyuan Zhang, Ruixuan Luo, Qi Su, Xu SunEMNLP 2022 · 9 citations
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan et al.NeurIPS 2022 · 132 citations
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 102 citations
