Particle Dual Averaging: Optimization of Mean Field Neural Network with Global Convergence Rate Analysis
Atsushi Nitanda, Denny Wu, Taiji Suzuki
Abstract
We propose the particle dual averaging (PDA) method, which generalizes the dual averaging method in convex optimization to the optimization over probability distributions with quantitative runtime guarantee. The algorithm consists of an inner loop and outer loop: the inner loop utilizes the Langevin algorithm to approximately solve for a stationary distribution, which is then optimized in the outer loop. The method can thus be interpreted as an extension of the Langevin algorithm to naturally handle nonlinear functional on the probability space. An important application of the proposed method is the optimization of neural network in the mean field regime, which is theoretically attractive due to the presence of nonlinear feature learning, but quantitative convergence rate can be challenging to obtain. By adapting finite-dimensional convex optimization theory into the space of measures, we analyze PDA in regularized empirical / expected risk minimization, and establish quantitative global convergence in learning two-layer mean field neural networks under more general settings. Our theoretical results are supported by numerical simulations on neural networks with reasonable size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 512e558c-f82b-4fa4-a23c-7d4d3f9fb11eCited by top-tier papers10
- Symmetric Mean-field Langevin Dynamics for Distributional Minimax ProblemsJuno Kim, Kakei Yamamoto, Kazusato Oko, Zhuoran Yang et al.ICLR 2024 · 14 citations
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 10 citations
- Global Convergence in Training Large-Scale TransformersCheng Gao, Yuan Cao, Zihao Li, Yihan He et al.NeurIPS 2024 · 10 citations
- Rethinking Information-theoretic Generalization: Loss Entropy Induced PAC BoundsYuxin Dong, Tieliang Gong, Hong Chen, Shujian Yu et al.ICLR 2024 · 8 citations
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 4 citations
Builds on8
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 104 citations
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
- Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov spaceTaiji Suzuki, Atsushi NitandaNeurIPS 2021 · 76 citations
Related papers
- Particle Stochastic Dual Coordinate Ascent: Exponential convergent algorithm for mean field neural network optimizationKazusato Oko, Taiji Suzuki, Atsushi Nitanda, Denny WuICLR 2022 · 8 citations
- Two-layer neural network on infinite dimensional data: global optimization guarantee in the mean-field regimeNaoki Nishikawa, Taiji Suzuki, Atsushi Nitanda, Denny WuNeurIPS 2022 · 7 citations
- Mean-field Underdamped Langevin Dynamics and its Spacetime DiscretizationQiang Fu, Ashia Camage WilsonICML 2024 · 5 citations
- Mean-Field Langevin Dynamics for Signed Measures via a Bilevel ApproachGuillaume Wang, Alireza Mousavi-Hosseini, Lénaïc ChizatNeurIPS 2024 · 8 citations
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 3 citations
