Understanding End-to-End Model-Based Reinforcement Learning Methods as Implicit Parameterization
Clement Gehring, Kenji Kawaguchi, Jiaoyang Huang, Leslie Pack Kaelbling
摘要
Estimating the per-state expected cumulative rewards is a critical aspect of reinforcement learning approaches, however the experience is obtained, but standard deep neural-network function-approximation methods are often inefficient in this setting. An alternative approach, exemplified by value iteration networks, is to learn transition and reward models of a latent Markov decision process whose value predictions fit the data. This approach has been shown empirically to converge faster to a more robust solution in many cases, but there has been little theoretical study of this phenomenon. In this paper, we explore such implicit representations of value functions via theory and focused experimentation. We prove that, for a linear parametrization, gradient descent converges to global optima despite nonlinearity and non-convexity introduced by the implicit representation. Furthermore, we derive convergence rates for both cases which allow us to identify conditions under which stochastic gradient descent (SGD) with this implicit representation converges substantially faster than its explicit counterpart. Finally, we provide empirical results in some simple domains that illustrate the theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The Benefits of Model-Based Generalization in Reinforcement LearningKenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen SchmidhuberICML 2023 · 被引用 18 次
- Model-Driven Deep Neural Network for Enhanced AoA Estimation Using 5G gNBShengheng Liu, Xingkang Li, Zihuan Mao, Peng Liu 等AAAI 2024 · 被引用 10 次
- Scaling up and Stabilizing Differentiable Planning with Implicit DifferentiationLinfeng Zhao, Huazhe Xu, Lawson L. S. WongICLR 2023 · 被引用 1 次
它引用的顶会 Paper5
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 被引用 87 次
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 被引用 47 次
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 被引用 47 次
相关 Paper
- Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence RateYufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran WangICML 2020 · 被引用 12 次
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- Provably Efficient Reinforcement Learning with Kernel and Neural Function ApproximationsZhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang 等NeurIPS 2020 · 被引用 48 次
- An Efficient Algorithm for Deep Stochastic Contextual BanditsTan Zhu, Guannan Liang, Chunjiang Zhu, Haining Li 等AAAI 2021 · 被引用 1 次
- Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningYuhui Wang, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio 等ICML 2025
