Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning
Younggyo Seo, Kimin Lee, Ignasi Clavera Gilaberte, Thanard Kurutach, Jinwoo Shin, Pieter Abbeel
摘要
Model-based reinforcement learning (RL) has shown great potential in various control tasks in terms of both sample-efficiency and final performance. However, learning a generalizable dynamics model robust to changes in dynamics remains a challenge since the target transition dynamics follow a multi-modal distribution. In this paper, we present a new model-based RL algorithm, coined trajectory-wise multiple choice learning, that learns a multi-headed dynamics model for dynamics generalization. The main idea is updating the most accurate prediction head to specialize each head in certain environments with similar dynamics, i.e., clustering environments. Moreover, we incorporate context learning, which encodes dynamicsspecific information from past experiences into the context latent vector, enabling the model to perform online adaptation to unseen environments. Finally, to utilize the specialized prediction heads more effectively, we propose an adaptive planning method, which selects the most accurate prediction head over a recent experience. Our method exhibits superior zero-shot generalization performance across a variety of control tasks, compared to state-of-the-art RL methods. Source code and videos are available at https://sites.google.com/view/trajectory-mcl .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline EnvironmentPhilip J. Ball, Cong Lu, Jack Parker-Holder, Stephen J. RobertsICML 2021 · 被引用 55 次
- Self-Paced Context Evaluation for Contextual Reinforcement LearningTheresa Eimer, André Biedenkapp, Frank Hutter, Marius LindauerICML 2021 · 被引用 35 次
- Dynamics Generalisation in Reinforcement Learning via Adaptive Context-Aware PoliciesMichael Beukman, Devon Jarvis, Richard Klein, Steven James 等NeurIPS 2023 · 被引用 29 次
- A Relational Intervention Approach for Unsupervised Dynamics Generalization in Model-Based Reinforcement LearningJiaxian Guo, Mingming Gong, Dacheng TaoICLR 2022 · 被引用 21 次
- DOMINO: Decomposed Mutual Information Optimization for Generalized Context in Meta-Reinforcement LearningYao Mu, Yuzheng Zhuang, Fei Ni, Bin Wang 等NeurIPS 2022 · 被引用 15 次
它引用的顶会 Paper4
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
- Dynamics-Aware EmbeddingsWilliam F. Whitney, Rajat Agarwal, Kyunghyun Cho, Abhinav GuptaICLR 2020
相关 Paper
- Learning and Planning Multi-Agent Tasks via an MoE-based World ModelZijie Zhao, Zhongyue Zhao, Kaixuan Xu, Yuqian Fu 等NeurIPS 2025 · 被引用 12 次
- Structure Detection for Contextual Reinforcement LearningTianyue Zhou, Jung-Hoon Cho, Cathy WuAAAI 2026
- Model-Based Transfer Learning for Contextual Reinforcement LearningJung-Hoon Cho, Vindula Jayawardana, Sirui Li, Cathy WuNeurIPS 2024 · 被引用 14 次
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-ScenariosXuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang 等AAAI 2026
- MetaDiffuser: Diffusion Model as Conditional Planner for Offline Meta-RLFei Ni, Jianye Hao, Yao Mu, Yifu Yuan 等ICML 2023 · 被引用 75 次
