On the Theoretical Properties of Noise Correlation in Stochastic Optimization
Aurélien Lucchi, Frank Proske, Antonio Orvieto, Francis R. Bach, Hans Kersting
Abstract
Studying the properties of stochastic noise to optimize complex non-convex functions has been an active area of research in the field of machine learning. Prior work [55, 50] has shown that the noise of stochastic gradient descent improves optimization by overcoming undesirable obstacles in the landscape. Moreover, injecting artificial Gaussian noise has become a popular idea to quickly escape saddle points. Indeed, in the absence of reliable gradient information, the noise is used to explore the landscape, but it is unclear what type of noise is optimal in terms of exploration ability. In order to narrow this gap in our knowledge, we study a general type of continuous-time non-Markovian process, based on fractional Brownian motion, that allows for the increments of the process to be correlated. This generalizes processes based on Brownian motion, such as the Ornstein-Uhlenbeck process. We demonstrate how to discretize such processes which gives rise to the new algorithm "fPGD". This method is a generalization of the known algorithms PGD and Anti-PGD [36] . We study the properties of fPGD both theoretically and empirically, demonstrating that it possesses exploration abilities that, in some cases, are favorable over PGD and Anti-PGD. These results open the field to novel ways to exploit noise for training machine learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 718b20e5-94ee-4533-b51e-bc7d46634ecaCited by top-tier papers3
- An SDE for Modeling SAM: Theory and InsightsEnea Monzio Compagnoni, Luca Biggio, Antonio Orvieto, Frank Norbert Proske et al.ICML 2023 · 25 citations
- Gradient Descent with Linearly Correlated Noise: Theory and Applications to Differential PrivacyAnastasia Koloskova, Ryan McKenna, Zachary Charles, John Keith Rush et al.NeurIPS 2023 · 24 citations
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li et al.CVPR 2026 · 1 citation
Builds on5
- The Heavy-Tail Phenomenon in SGDMert Gürbüzbalaban, Umut Simsekli, Lingjiong ZhuICML 2021 · 165 citations
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 155 citations
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 122 citations
- Anticorrelated Noise Injection for Improved GeneralizationAntonio Orvieto, Hans Kersting, Frank Proske, Francis R. Bach et al.ICML 2022 · 58 citations
Related papers
- Learning Fractional White Noises in Neural Stochastic Differential EquationsAnh Tong, Thanh Nguyen-Tang, Toan M. Tran, Jaesik ChoiNeurIPS 2022 · 17 citations
- Fractional Langevin Dynamics for Combinatorial Optimization via Polynomial-Time EscapeShiyue Wang, Ziao Guo, Changhong Lu, Junchi YanNeurIPS 2025 · 5 citations
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 38 citations
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 58 citations
- Pink Noise Is All You Need: Colored Noise Exploration in Deep Reinforcement LearningOnno Eberhard, Jakob J. Hollenstein, Cristina Pinneri, Georg MartiusICLR 2023
