Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
Yufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen, Zhaoran Wang
Abstract
Temporal-difference and Q-learning play a key role in deep reinforcement learning, where they are empowered by expressive nonlinear function approximators such as neural networks. At the core of their empirical successes is the learned feature representation, which embeds rich observations, e.g., images and texts, into the latent space that encodes semantic structures. Meanwhile, the evolution of such a feature representation is crucial to the convergence of temporal-difference and Q-learning. In particular, temporal-difference learning converges when the function approximator is linear in a feature representation, which is fixed throughout learning, and possibly diverges otherwise. We aim to answer the following questions: When the function approximator is a neural network, how does the associated feature representation evolve? If it converges, does it converge to the optimal one? We prove that, utilizing an overparameterized two-layer neural network, temporal-difference and Q-learning globally minimize the mean-squared projected Bellman error at a sublinear rate. Moreover, the associated feature representation converges to the optimal one, generalizing the previous analysis of Cai et al. (2019) in the neural tangent kernel regime, where the associated feature representation stabilizes at the initial one. The key to our analysis is a mean-field perspective, which connects the evolution of a finite-dimensional parameter to its limiting counterpart over an infinite-dimensional Wasserstein space. Our analysis generalizes to soft Q-learning, which is further connected to policy gradient.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 270c506b-0c70-4352-bc51-46ff4f48449aCited by top-tier papers7
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 155 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
- Offline RL with Observation Histories: Analyzing and Improving Sample ComplexityJoey Hong, Anca D. Dragan, Sergey LevineICLR 2024 · 8 citations
- Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regimeAndrea Agazzi, Jianfeng LuICLR 2021 · 3 citations
- Towards a better understanding of representation dynamics under TD-learningYunhao Tang, Rémi MunosICML 2023 · 3 citations
Builds on3
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 18 citations
Related papers
- Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-CriticYufeng Zhang, Siyu Chen, Zhuoran Yang, Michael I. Jordan et al.NeurIPS 2021 · 6 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
- On the Performance of Temporal Difference Learning With Neural NetworksHaoxing Tian, Ioannis Ch. Paschalidis, Alex OlshevskyICLR 2023
- Mean Field Langevin Actor-Critic: Faster Convergence and Global Optimality beyond Lazy LearningKakei Yamamoto, Kazusato Oko, Zhuoran Yang, Taiji SuzukiICML 2024 · 2 citations
- Provably Efficient Neural GTD for Off-Policy LearningHoi-To Wai, Zhuoran Yang, Zhaoran Wang, Mingyi HongNeurIPS 2020 · 7 citations
