Particle Stochastic Dual Coordinate Ascent: Exponential convergent algorithm for mean field neural network optimization
Kazusato Oko, Taiji Suzuki, Atsushi Nitanda, Denny Wu
Abstract
We introduce Particle-SDCA, a gradient-based optimization algorithm for two-layer neural networks in the mean field regime that achieves exponential convergence rate in regularized empirical risk minimization. The proposed algorithm can be regarded as an infinite dimensional extension of Stochastic Dual Coordinate Ascent (SDCA) in the probability space: we exploit the convexity of the dual problem, for which the coordinate-wise proximal gradient method can be applied. Our proposed method inherits advantages of the original SDCA, including (i) exponential convergence (with respect to the outer iteration steps), and (ii) better dependency on the sample size and condition number than the full-batch gradient method. One technical challenge in implementing the SDCA update is the intractable integral over the entire parameter space at every step. To overcome this limitation, we propose a tractable particle method that approximately solves the dual problem, and an importance re-weighted technique to reduce the computational cost. The convergence rate of our method is verified by numerical experiments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get baf6e71c-b82d-4813-89f7-3b073707eda1Cited by top-tier papers7
- Improved Particle Approximation Error for Mean Field Neural NetworksAtsushi NitandaNeurIPS 2024 · 18 citations
- Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context LearningDake Bu, Wei Huang, Andi Han, Atsushi Nitanda et al.NeurIPS 2024 · 11 citations
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 10 citations
- Two-layer neural network on infinite dimensional data: global optimization guarantee in the mean-field regimeNaoki Nishikawa, Taiji Suzuki, Atsushi Nitanda, Denny WuNeurIPS 2022 · 7 citations
- Primal and Dual Analysis of Entropic Fictitious Play for Finite-sum ProblemsAtsushi Nitanda, Kazusato Oko, Denny Wu, Nobuhito Takenouchi et al.ICML 2023 · 4 citations
Related papers
- Particle Dual Averaging: Optimization of Mean Field Neural Network with Global Convergence Rate AnalysisAtsushi Nitanda, Denny Wu, Taiji SuzukiNeurIPS 2021 · 32 citations
- Quantitative Propagation of Chaos for SGD in Wide Neural NetworksValentin De Bortoli, Alain Durmus, Xavier Fontaine, Umut SimsekliNeurIPS 2020 · 36 citations
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 4 citations
- Uniform-in-time propagation of chaos for the mean-field gradient Langevin dynamicsTaiji Suzuki, Atsushi Nitanda, Denny WuICLR 2023
- Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model EnsembleAtsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai, Mizuki Sakaguchi et al.ICML 2025
