Improved high-dimensional estimation with Langevin dynamics and stochastic weight averaging
Stanley Wei, Alex Damian, Jason D. Lee
摘要
Significant recent work has studied the ability of gradient descent to recover a hidden planted direction in different high-dimensional settings, including tensor PCA and single-index models. The key quantity that governs the ability of gradient descent to traverse these landscapes is the information exponent (Ben Arous et al., (2021)), which corresponds to the order of the saddle at initialization in the population landscape. Ben Arous et al., (2021) showed that samples were necessary and sufficient for online SGD to recover , and Ben Arous et al., (2020) proved a similar lower bound for Langevin dynamics. More recently, Damian et al., (2023) showed it was possible to circumvent these lower bounds by running gradient descent on a smoothed landscape, and that this algorithm succeeds with samples, which is optimal in the worst case. This raises the question of whether it is possible to achieve the same rate without explicit smoothing. In this paper, we show that Langevin dynamics can succeed with samples if one considers the average iterate, rather than the last iterate. The key idea is that the combination of noise-injection and iterate averaging is able to emulate the effect of landscape smoothing. We apply this result to both the tensor PCA and single-index model settings. Finally, we conjecture that minibatch SGD can also achieve the same rate without adding any additional noise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Learning single-index models with shallow neural networksAlberto Bietti, Joan Bruna, Clayton Sanford, Min Jae SongNeurIPS 2022 · 被引用 119 次
- Phase retrieval in high dimensions: Statistical and computational phase transitionsAntoine Maillard, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2020 · 被引用 73 次
- Neural network learns low-dimensional polynomials with SGD near the information-theoretic limitJason D. Lee, Kazusato Oko, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 49 次
- Learning single index models via harmonic decompositionNirmit Joshi, Hugo Koubbi, Theodor Misiakiewicz, Nati SrebroNeurIPS 2025 · 被引用 8 次
- Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent AnalysisYunwei Ren, Jason D. LeeNeurIPS 2025 · 被引用 7 次
相关 Paper
- Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index ModelsAlex Damian, Eshaan Nichani, Rong Ge, Jason D. LeeNeurIPS 2023 · 被引用 67 次
- On Single-Index Models beyond Gaussian DataAaron Zweig, Loucas Pillaud-Vivien, Joan BrunaNeurIPS 2023 · 被引用 17 次
- Gradient-Based Feature Learning under Structured DataAlireza Mousavi-Hosseini, Denny Wu, Taiji Suzuki, Murat A. ErdogduNeurIPS 2023 · 被引用 36 次
- High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizesAukosh Jagannath, Taj Jones-McCormick, Varnan SarangianICLR 2026 · 被引用 1 次
- Symmetric Single Index LearningAaron Zweig, Joan BrunaICLR 2024 · 被引用 4 次
