Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
Zeke Xie, Li Yuan, Zhanxing Zhu, Masashi Sugiyama
摘要
It is well-known that stochastic gradient noise (SGN) acts as implicit regularization for deep learning and is essentially important for both optimization and generalization of deep networks. Some works attempted to artificially simulate SGN by injecting random noise to improve deep learning. However, it turned out that the injected simple random noise cannot work as well as SGN, which is anisotropic and parameter-dependent. For simulating SGN at low computational costs and without changing the learning rate or batch size, we propose the Positive-Negative Momentum (PNM) approach that is a powerful alternative to conventional Momentum in classic optimizers. The introduced PNM method maintains two approximate independent momentum terms. Then, we can control the magnitude of SGN explicitly by adjusting the momentum difference. We theoretically prove the convergence guarantee and the generalization advantage of PNM over Stochastic Gradient Descent (SGD). By incorporating PNM into the two conventional optimizers, SGD with Momentum and Adam, our extensive experiments empirically verified the significant advantage of the PNM-based variants over the corresponding conventional Momentum-based optimizers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui 等ICLR 2022 · 被引用 213 次
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato 等ICML 2022 · 被引用 65 次
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato 等NeurIPS 2023 · 被引用 38 次
- Sparse Double Descent: Where Network Pruning Aggravates OverfittingZheng He, Zeke Xie, Quanzhi Zhu, Zengchang QinICML 2022 · 被引用 36 次
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and ImplicationsZiqiao Wang, Yongyi MaoICLR 2022 · 被引用 33 次
它引用的顶会 Paper3
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 被引用 165 次
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
相关 Paper
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- The Marginal Value of Momentum for Small Learning Rate SGDRunzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu 等ICLR 2024 · 被引用 14 次
- Towards understanding how momentum improves generalization in deep learningSamy Jelassi, Yuanzhi LiICML 2022 · 被引用 53 次
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun 等NeurIPS 2022 · 被引用 29 次
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li 等CVPR 2026 · 被引用 1 次
