Generalized Hidden Parameter MDPs: Transferable Model-Based RL in a Handful of Trials
Christian F. Perez, Felipe Petroski Such, Theofanis Karaletsos
摘要
There is broad interest in creating RL agents that can solve many (related) tasks and adapt to new tasks and environments after initial training. Model-based RL leverages learned surrogate models that describe dynamics and rewards of individual tasks, such that planning in a good surrogate can lead to good control of the true system. Rather than solving each task individually from scratch, hierarchical models can exploit the fact that tasks are often related by (unobserved) causal factors of variation in order to achieve efficient generalization, as in learning how the mass of an item affects the force required to lift it can generalize to previously unobserved masses. We propose Generalized Hidden Parameter MDPs (GHP-MDPs) that describe a family of MDPs where both dynamics and reward can change as a function of hidden parameters that vary across tasks. The GHP-MDP augments model-based RL with latent variables that capture these hidden parameters, facilitating transfer across tasks. We also explore a variant of the model that incorporates explicit latent structure mirroring the causal factors of variation across tasks (for instance: agent properties, environmental factors, and goals). We experimentally demonstrate state-of-the-art performance and sample-efficiency on a new challenging MuJoCo task using reward and dynamics latent spaces, while beating a previous state-of-the-art baseline with > 10× less data. Using test-time inference of the latent variables, our approach generalizes in a single episode to novel combinations of dynamics and reward, and to novel rewards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Causal Curiosity: RL Agents Discovering Self-supervised Experiments for Causal Representation LearningSumedh A. Sontakke, Arash Mehrjou, Laurent Itti, Bernhard SchölkopfICML 2021 · 被引用 73 次
- Contrastive Behavioral Similarity Embeddings for Generalization in Reinforcement LearningRishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, Marc G. BellemareICLR 2021 · 被引用 27 次
- MAMBA: an Effective World Model Approach for Meta-Reinforcement LearningZohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler 等ICLR 2024 · 被引用 15 次
- Zero-Shot Assistance in Sequential Decision ProblemsSebastiaan De Peuter, Samuel KaskiAAAI 2023 · 被引用 6 次
- Learning Robust State Abstractions for Hidden-Parameter Block MDPsAmy Zhang, Shagun Sodhani, Khimya Khetarpal, Joelle PineauICLR 2021 · 被引用 5 次
相关 Paper
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
- Improving Generalization in Meta-RL with Imaginary Tasks from Latent Dynamics MixtureSuyoung Lee, Sae-Young ChungNeurIPS 2021 · 被引用 23 次
- Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement LearningXinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng LvKDD 2026
- Single Episode Policy Transfer in Reinforcement LearningJiachen Yang, Brenden K. Petersen, Hongyuan Zha, Daniel M. FaissolICLR 2020 · 被引用 38 次
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos 等ICML 2020 · 被引用 153 次
