Particle Dual Averaging: Optimization of Mean Field Neural Network with Global Convergence Rate Analysis
Atsushi Nitanda, Denny Wu, Taiji Suzuki
摘要
We propose the particle dual averaging (PDA) method, which generalizes the dual averaging method in convex optimization to the optimization over probability distributions with quantitative runtime guarantee. The algorithm consists of an inner loop and outer loop: the inner loop utilizes the Langevin algorithm to approximately solve for a stationary distribution, which is then optimized in the outer loop. The method can thus be interpreted as an extension of the Langevin algorithm to naturally handle nonlinear functional on the probability space. An important application of the proposed method is the optimization of neural network in the mean field regime, which is theoretically attractive due to the presence of nonlinear feature learning, but quantitative convergence rate can be challenging to obtain. By adapting finite-dimensional convex optimization theory into the space of measures, we analyze PDA in regularized empirical / expected risk minimization, and establish quantitative global convergence in learning two-layer mean field neural networks under more general settings. Our theoretical results are supported by numerical simulations on neural networks with reasonable size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Symmetric Mean-field Langevin Dynamics for Distributional Minimax ProblemsJuno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang 等ICLR 2024 · 被引用 14 次
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 被引用 10 次
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He 等NeurIPS 2024 · 被引用 10 次
- Rethinking Information-theoretic Generalization: Loss Entropy Induced PAC BoundsYuxin Dong, Tieliang Gong, Hong Chen, Shujian Yu 等ICLR 2024 · 被引用 8 次
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 被引用 4 次
它引用的顶会 Paper8
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 被引用 128 次
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 被引用 104 次
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 被引用 82 次
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 被引用 76 次
相关 Paper
- Particle Stochastic Dual Coordinate Ascent: Exponential convergent algorithm for mean field neural network optimizationKazusato Oko, Taiji Suzuki, Atsushi Nitanda, Denny WuICLR 2022 · 被引用 8 次
- Two-layer neural network on infinite dimensional data: global optimization guarantee in the mean-field regimeNaoki Nishikawa, Taiji Suzuki, Atsushi Nitanda, Denny WuNeurIPS 2022 · 被引用 7 次
- Mean-field Underdamped Langevin Dynamics and its Spacetime DiscretizationQiang Fu, Ashia Camage WilsonICML 2024 · 被引用 5 次
- Mean-Field Langevin Dynamics for Signed Measures via a Bilevel ApproachGuillaume Wang, Alireza Mousavi-Hosseini, Lénaïc ChizatNeurIPS 2024 · 被引用 8 次
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 被引用 3 次
