Model-Augmented Actor-Critic: Backpropagating through Paths
Ignasi Clavera, Yao Fu, Pieter Abbeel
Abstract
Current model-based reinforcement learning approaches use the model simply as a learned black-box simulator to augment the data for policy optimization or value function learning. In this paper, we show how to make more effective use of the model by exploiting its differentiability. We construct a policy optimization algorithm that uses the pathwise derivative of the learned model and policy across future timesteps. Instabilities of learning across many timesteps are prevented by using a terminal value function, learning the policy in an actor-critic fashion. Furthermore, we present a derivation on the monotonic improvement of our objective in terms of the gradient error in the model and value function. We show that our approach (i) is consistently more sample efficient than existing state-of-the-art model-based algorithms, (ii) matches the asymptotic performance of model-free algorithms, and (iii) scales to long horizons, a regime where typically past model-based approaches have struggled.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b93057c-6203-4cef-93b9-99c212c14f9dCited by top-tier papers40
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
- RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2022 · 168 citations
- Accelerated Policy Learning with Parallel Differentiable SimulationJie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos et al.ICLR 2022 · 141 citations
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 120 citations
Builds on1
Related papers
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu et al.ICLR 2022 · 10 citations
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 47 citations
- Deep Model-Based Reinforcement Learning via Estimated Uncertainty and Conservative Policy OptimizationQi Zhou, Houqiang Li, Jie WangAAAI 2020 · 17 citations
- Trust the Model When It Is Confident: Masked Model-based Actor-CriticFeiyang Pan, Jia He, Dandan Tu, Qing HeNeurIPS 2020 · 65 citations
- On-Policy Model Errors in Reinforcement LearningLukas P. Fröhlich, Maksym Lefarov, Melanie N. Zeilinger, Felix BerkenkampICLR 2022 · 6 citations
