Improved statistical and computational complexity of the mean-field Langevin dynamics under structured data
Atsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny Wu
摘要
The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks via (noisy) gradient descent. Recent works have shown that MFLD globally minimizes an entropy-regularized convex functional in the space of measures. However, all prior analyses assumed the infinite-particle or continuous-time limit, and cannot handle stochastic gradient updates. We provide an general framework to prove a uniform-in-time propagation of chaos for MFLD that takes into account the errors due to finite-particle approximation, timediscretization, and stochastic gradient approximation. To demonstrate the wide applicability of this framework, we establish quantitative convergence rate guarantees to the regularized global optimal solution under (i) a wide range of learning problems such as neural network in the meanfield regime and MMD minimization, and (ii) different gradient estimators including SGD and SVRG. Despite the generality of our results, we achieve an improved convergence rate in both the SGD and SVRG settings when specialized to the standard Langevin dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model EnsembleAtsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai, Mizuki Sakaguchi 等ICML 2025
- Learning Multi-Index Models with Neural Networks via Mean-Field Langevin DynamicsAlireza Mousavi-Hosseini, Denny Wu, Murat A. ErdogduICLR 2025
- Uniform-in-time propagation of chaos for the mean-field gradient Langevin dynamicsTaiji Suzuki, Atsushi Nitanda, Denny WuICLR 2023
- Robust Feature Learning for Multi-Index Models in High DimensionsAlireza Mousavi-Hosseini, Adel Javanmard, Murat A. ErdogduICLR 2025
它引用的顶会 Paper7
- Kernel Stein Discrepancy DescentAnna Korba, Pierre-Cyril Aubin-Frankowski, Szymon Majewski, Pierre AblinICML 2021 · 被引用 64 次
- Quantitative Propagation of Chaos for SGD in Wide Neural NetworksValentin De Bortoli, Alain Durmus, Xavier Fontaine, Umut SimsekliNeurIPS 2020 · 被引用 36 次
- A Dynamical Central Limit Theorem for Shallow Neural NetworksZhengdao Chen, Grant M. Rotskoff, Joan Bruna, Eric Vanden-EijndenNeurIPS 2020 · 被引用 33 次
- Particle Dual Averaging: Optimization of Mean Field Neural Network with Global Convergence Rate AnalysisAtsushi Nitanda, Denny Wu, Taiji SuzukiNeurIPS 2021 · 被引用 32 次
- Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to OptimizationYuri Kinoshita, Taiji SuzukiNeurIPS 2022 · 被引用 24 次
相关 Paper
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 被引用 10 次
- Improved Particle Approximation Error for Mean Field Neural NetworksAtsushi NitandaNeurIPS 2024 · 被引用 18 次
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 被引用 3 次
- Mean-field Underdamped Langevin Dynamics and its Spacetime DiscretizationQiang Fu, Ashia Camage WilsonICML 2024 · 被引用 5 次
- Symmetric Mean-field Langevin Dynamics for Distributional Minimax ProblemsJuno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang 等ICLR 2024 · 被引用 14 次
