How to Escape Sharp Minima with Random Perturbations
Kwangjun Ahn, Ali Jadbabaie, Suvrit Sra
摘要
Modern machine learning applications have witnessed the remarkable success of optimization algorithms that are designed to find flat minima. Motivated by this design choice, we undertake a formal study that (i) formulates the notion of flat minima, and (ii) studies the complexity of finding them. Specifically, we adopt the trace of the Hessian of the cost function as a measure of flatness, and use it to formally define the notion of approximate flat minima. Under this notion, we then analyze algorithms that find approximate flat minima efficiently. For general cost functions, we discuss a gradient-based algorithm that finds an approximate flat local minimum efficiently. The main component of the algorithm is to use gradients computed from randomly perturbed iterates to estimate a direction that leads to flatter minima. For the setting where the cost function is an empirical risk over training data, we present a faster algorithm that is inspired by a recently proposed practical algorithm called sharpness-aware minimization, supporting its success in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Explicit Eigenvalue Regularization Improves Sharpness-Aware MinimizationHaocheng Luo, Tuan Truong, Tung Pham, Mehrtash Harandi 等NeurIPS 2024 · 被引用 28 次
- Fundamental Convergence Analysis of Sharpness-Aware MinimizationPham Duy Khanh, Hoang-Chau Luong, Boris S. Mordukhovich, Dat Ba TranNeurIPS 2024 · 被引用 27 次
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等NeurIPS 2025 · 被引用 8 次
- Flexible Sharpness-Aware Personalized Federated LearningXinda Xing, Qiugang Zhan, Xiurui Xie, Yuning Yang 等AAAI 2025 · 被引用 5 次
- Gradient Extrapolation for Debiased Representation LearningIhab Asaad, Maha Shadaydeh, Joachim DenzlerICCV 2025 · 被引用 4 次
它引用的顶会 Paper24
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 被引用 165 次
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
相关 Paper
- Flat Minima and Generalization: Insights from Stochastic Convex OptimizationMatan Schliserman, Shira Vansover-Hager, Tomer KorenICML 2026 · 被引用 2 次
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou 等CVPR 2023
- Entropic gradient descent algorithms and wide flat minimaFabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Gabriele Perugini 等ICLR 2021 · 被引用 38 次
- An Adaptive Policy to Employ Sharpness-Aware MinimizationWeisen Jiang, Hansi Yang, Yu Zhang, James T. KwokICLR 2023 · 被引用 2 次
- Improving Sharpness-Aware Minimization by LookaheadRunsheng Yu, Youzhi Zhang, James T. KwokICML 2024 · 被引用 1 次
