Uniform-in-time propagation of chaos for the mean-field gradient Langevin dynamics
Taiji Suzuki, Atsushi Nitanda, Denny Wu
摘要
Mean-field Langevin dynamics (MFLD) is an optimization method derived by taking the mean-field limit of noisy gradient descent for two-layer neural networks in the mean-field regime. Recently, the propagation of chaos (PoC) for MFLD has gained attention as it provides a quantitative characterization of the optimization complexity in terms of the number of particles and iterations. A remarkable progress by Chen et al. ( 2022 ) showed that the approximation error due to finite particles remains uniform in time and diminishes as the number of particles increases. In this paper, by refining the defective log-Sobolev inequalitya key result from that earlier work-under the neural network training setting, we establish an improved PoC result for MFLD, which removes the exponential dependence on the regularization coefficient from the particle approximation term of the optimization complexity. As an application, we propose a PoC-based model ensemble strategy with theoretical guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Improved Particle Approximation Error for Mean Field Neural NetworksAtsushi NitandaNeurIPS 2024 · 被引用 18 次
- Feature learning via mean-field Langevin dynamics: classifying sparse parities and beyondTaiji Suzuki, Denny Wu, Kazusato Oko, Atsushi NitandaNeurIPS 2023 · 被引用 17 次
- Mean-field Analysis on Two-layer Neural Networks from a Kernel PerspectiveShokichi Takakura, Taiji SuzukiICML 2024 · 被引用 12 次
- Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reductionTaiji Suzuki, Denny Wu, Atsushi NitandaNeurIPS 2023 · 被引用 10 次
- Learning of Population Dynamics: Inverse Optimization Meets JKO SchemeMikhail Persiianov, Jiawei Chen, Petr Mokrov, Alexander Tyurin 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper14
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 被引用 654 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 被引用 272 次
相关 Paper
- Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model EnsembleAtsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai, Mizuki Sakaguchi 等ICML 2025
- Improved statistical and computational complexity of the mean-field Langevin dynamics under structured dataAtsushi Nitanda, Kazusato Oko, Taiji Suzuki, Denny WuICLR 2024 · 被引用 4 次
- Mirror Mean-Field Langevin DynamicsAnming Gu, Juno KimICML 2026 · 被引用 3 次
- Quantitative Propagation of Chaos for SGD in Wide Neural NetworksValentin De Bortoli, Alain Durmus, Xavier Fontaine, Umut SimsekliNeurIPS 2020 · 被引用 36 次
- Global Convergence of Three-layer Neural Networks in the Mean Field RegimeHuy Tuan Pham, Phan-Minh NguyenICLR 2021 · 被引用 23 次
