Continuous MDP Homomorphisms and Homomorphic Policy Gradient
Sahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger, Doina Precup
Abstract
Abstraction has been widely studied as a way to improve the efficiency and generalization of reinforcement learning algorithms. In this paper, we study abstraction in the continuous-control setting. We extend the definition of MDP homomorphisms to encompass continuous actions in continuous state spaces. We derive a policy gradient theorem on the abstract MDP, which allows us to leverage approximate symmetries of the environment for policy optimization. Based on this theorem, we propose an actor-critic algorithm that is able to learn the policy and the MDP homomorphism map simultaneously, using the lax bisimulation metric. We demonstrate the effectiveness of our method on benchmark tasks in the DeepMind Control Suite. Our method's ability to utilize MDP homomorphisms for representation learning leads to improved performance when learning from pixel observations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d97dd8f-f707-4595-b058-dcc1a27bca88Cited by top-tier papers13
- For SALE: State-Action Representation Learning for Deep Reinforcement LearningScott Fujimoto, Wei-Di Chang, Edward J. Smith, Shixiang Gu et al.NeurIPS 2023 · 128 citations
- EDGI: Equivariant Diffusion for Planning with Embodied AgentsJohann Brehmer, Joey Bose, Pim de Haan, Taco S. CohenNeurIPS 2023 · 52 citations
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu et al.ICML 2024 · 30 citations
- E(3)-Equivariant Actor-Critic Methods for Cooperative Multi-Agent Reinforcement LearningDingyang Chen, Qi ZhangICML 2024 · 10 citations
- Reinforcement Learning with Euclidean Data Augmentation for State-Based Continuous ControlJinzhu Luo, Dingyang Chen, Qi ZhangNeurIPS 2024 · 5 citations
Builds on21
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
Related papers
- Building Minimal and Reusable Causal State Abstractions for Reinforcement LearningZizhao Wang, Caroline Wang, Xuesu Xiao, Yuke Zhu et al.AAAI 2024 · 9 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
- Towards Robust Bisimulation Metric LearningMete Kemertas, Tristan Aumentado-ArmstrongNeurIPS 2021 · 68 citations
- Bisimulation Metric for Model Predictive ControlYutaka Shimizu, Masayoshi TomizukaICLR 2025
- Bisimulation Makes Analogies in Goal-Conditioned Reinforcement LearningPhilippe Hansen-Estruch, Amy Zhang, Ashvin Nair, Patrick Yin et al.ICML 2022 · 39 citations
