Mean Field Langevin Actor-Critic: Faster Convergence and Global Optimality beyond Lazy Learning
Kakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji Suzuki
摘要
This work explores the feature learning capabilities of deep reinforcement learning algorithms in the pursuit of optimal policy determination. We particularly examine an over-parameterized neural actor-critic framework within the meanfield regime, where both actor and critic components undergo updates via policy gradient and temporal-difference (TD) learning, respectively. We introduce the mean-field Langevin TD learning (MFLTD) method, enhancing mean-field Langevin dynamics with proximal TD updates for critic policy evaluation, and assess its performance against conventional approaches through numerical analysis. Additionally, for actor policy updates, we present the mean-field Langevin policy gradient (MFLPG), employing policy gradient techniques augmented by Wasserstein gradient flows for parameter space exploration. Our findings demonstrate that MFLTD accurately identifies the true value function, while MFLPG ensures linear convergence of actor sequences towards the globally optimal policy, considering a Kullback-Leibler divergence regularized framework. Through both time particle and discretized analysis, we substantiate the linear convergence guarantees of our neural actor-critic algorithms, representing a notable contribution to neural reinforcement learning focusing on global optimality and feature learning, extending the existing understanding beyond the conventional scope of lazy training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Non-convex entropic mean-field optimization via Best Response flowRazvan-Andrei Lascu, Mateusz B. MajkaNeurIPS 2025 · 被引用 3 次
- From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous EnvironmentsSaket Tiwari, Tejas Kotwal, George Dimitri KonidarisICLR 2026
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 被引用 82 次
- Impact of Representation Learning in Linear BanditsJiaqi Yang, Wei Hu, Jason D. Lee, Simon Shaolei DuICLR 2021 · 被引用 58 次
相关 Paper
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo 等ICML 2025
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field TheoryYufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen 等NeurIPS 2020 · 被引用 12 次
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-CriticYufeng Zhang, Siyu Chen, Zhuoran Yang, Michael I. Jordan 等NeurIPS 2021 · 被引用 6 次
- Single-Timescale Actor-Critic Provably Finds Globally Optimal PolicyZuyue Fu, Zhuoran Yang, Zhaoran WangICLR 2021 · 被引用 52 次
- Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte CarloHaque Ishfaq, Qingfeng Lan, Pan Xu, A. Rupam Mahmood 等ICLR 2024 · 被引用 33 次
