Learning Robust State Abstractions for Hidden-Parameter Block MDPs
Amy Zhang, Shagun Sodhani, Khimya Khetarpal, Joelle Pineau
Abstract
Many control tasks exhibit similar dynamics that can be modeled as having common latent structure. Hidden-Parameter Markov Decision Processes (HiP-MDPs) explicitly model this structure to improve sample efficiency in multi-task settings. However, this setting makes strong assumptions on the observability of the state that limit its application in real-world scenarios with rich observation spaces. In this work, we leverage ideas of common structure from the HiP-MDP setting, and extend it to enable robust state abstractions inspired by Block MDPs. We derive instantiations of this new framework for both multi-task reinforcement learning (MTRL) and meta-reinforcement learning (Meta-RL) settings. Further, we provide transfer and generalization bounds based on task and state similarity, along with sample complexity bounds that depend on the aggregate number of samples across tasks, rather than the number of tasks, a significant improvement over prior work that use the same environment assumptions. To further demonstrate the efficacy of the proposed method, we empirically compare and show improvement over multi-task and meta-reinforcement learning baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3d937d0-8cb1-467e-a6b4-d5a3ee7c1637Cited by top-tier papers16
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement LearningBiwei Huang, Fan Feng, Chaochao Lu, Sara Magliacane et al.ICLR 2022 · 75 citations
- Efficient Reinforcement Learning in Block MDPs: A Model-free Representation Learning approachXuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang et al.ICML 2022 · 65 citations
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentPhilip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. RobertsICML 2021 · 55 citations
- Cross-Trajectory Representation Learning for Zero-Shot Generalization in RLBogdan Mazoure, Ahmed M. Ahmed, R. Devon Hjelm, Andrey Kolobov et al.ICLR 2022 · 30 citations
Builds on8
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos et al.AAAI 2021 · 506 citations
- Meta-Learning without MemorizationMingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine et al.ICLR 2020 · 201 citations
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos et al.ICML 2020 · 153 citations
- Sharing Knowledge in Multi-Task Deep Reinforcement LearningCarlo D'Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli et al.ICLR 2020 · 148 citations
Related papers
- Generalized Hidden Parameter MDPs: Transferable Model-Based RL in a Handful of TrialsChristian F. Perez, Felipe Petroski Such, Theofanis KaraletsosAAAI 2020 · 39 citations
- Provable Benefits of Multi-task RL under Non-Markovian Decision Making ProcessesRuiquan Huang, Yuan Cheng, Jing Yang, Vincent Tan et al.ICLR 2024
- MAMBA: an Effective World Model Approach for Meta-Reinforcement LearningZohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler et al.ICLR 2024 · 15 citations
- When Is Generalizable Reinforcement Learning Tractable?Dhruv Malik, Yuanzhi Li, Pradeep RavikumarNeurIPS 2021 · 32 citations
- Structure Detection for Contextual Reinforcement LearningTianyue Zhou, Jung-Hoon Cho, Cathy WuAAAI 2026
