Understanding End-to-End Model-Based Reinforcement Learning Methods as Implicit Parameterization
Clement Gehring, Kenji Kawaguchi, Jiaoyang Huang, Leslie Pack Kaelbling
Abstract
Estimating the per-state expected cumulative rewards is a critical aspect of reinforcement learning approaches, however the experience is obtained, but standard deep neural-network function-approximation methods are often inefficient in this setting. An alternative approach, exemplified by value iteration networks, is to learn transition and reward models of a latent Markov decision process whose value predictions fit the data. This approach has been shown empirically to converge faster to a more robust solution in many cases, but there has been little theoretical study of this phenomenon. In this paper, we explore such implicit representations of value functions via theory and focused experimentation. We prove that, for a linear parametrization, gradient descent converges to global optima despite nonlinearity and non-convexity introduced by the implicit representation. Furthermore, we derive convergence rates for both cases which allow us to identify conditions under which stochastic gradient descent (SGD) with this implicit representation converges substantially faster than its explicit counterpart. Finally, we provide empirical results in some simple domains that illustrate the theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0cd49db-e8bc-4282-b54f-509a8bbbe1ebCited by top-tier papers3
- The Benefits of Model-Based Generalization in Reinforcement LearningKenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen SchmidhuberICML 2023 · 18 citations
- Model-Driven Deep Neural Network for Enhanced AoA Estimation Using 5G gNBShengheng Liu, Xingkang Li, Zihuan Mao, Peng Liu et al.AAAI 2024 · 10 citations
- Scaling up and Stabilizing Differentiable Planning with Implicit DifferentiationLinfeng Zhao, Huazhe Xu, Lawson L. S. WongICLR 2023 · 1 citation
Builds on5
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More DepthKeyulu Xu, Mozhi Zhang, Stefanie Jegelka, Kenji KawaguchiICML 2021 · 87 citations
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 47 citations
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit LayersKenji KawaguchiICLR 2021 · 47 citations
Related papers
- Generative Adversarial Imitation Learning with Neural Network Parameterization: Global Optimality and Convergence RateYufeng Zhang, Qi Cai, Zhuoran Yang, Zhaoran WangICML 2020 · 12 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
- Provably Efficient Reinforcement Learning with Kernel and Neural Function ApproximationsZhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang et al.NeurIPS 2020 · 48 citations
- An Efficient Algorithm for Deep Stochastic Contextual BanditsTan Zhu, Guannan Liang, Chunjiang Zhu, Haining Li et al.AAAI 2021 · 1 citation
- Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term PlanningYuhui Wang, Qingyuan Wu, Dylan R. Ashley, Francesco Faccio et al.ICML 2025
