One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RL
Saurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea Finn
摘要
While reinforcement learning algorithms can learn effective policies for complex tasks, these policies are often brittle to even minor task variations, especially when variations are not explicitly provided during training. One natural approach to this problem is to train agents with manually specified variation in the training task or environment. However, this may be infeasible in practical situations, either because making perturbations is not possible, or because it is unclear how to choose suitable perturbation strategies without sacrificing performance. The key insight of this work is that learning diverse behaviors for accomplishing a task can directly lead to behavior that generalizes to varying environments, without needing to perform explicit perturbations during training. By identifying multiple solutions for the task in a single environment during training, our approach can generalize to new situations by abandoning solutions that are no longer effective and adopting those that are. We theoretically characterize a robustness set of environments that arises from our algorithm and empirically find that our diversity-driven approach can extrapolate to various changes in the environment and task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina 等ICLR 2024 · 被引用 332 次
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 被引用 157 次
- Learning more skills through optimistic explorationDJ Strouse, Kate Baumli, David Warde-Farley, Volodymyr Mnih 等ICLR 2022 · 被引用 53 次
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 被引用 48 次
- Towards Diverse Behaviors: A Benchmark for Imitation Learning with Human DemonstrationsXiaogang Jia, Denis Blessing, Xinkai Jiang, Moritz Reuss 等ICLR 2024 · 被引用 48 次
它引用的顶会 Paper5
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar 等ICLR 2020 · 被引用 475 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
相关 Paper
- Robust Policy Learning over Multiple Uncertainty SetsAnnie Xie, Shagun Sodhani, Chelsea Finn, Joelle Pineau 等ICML 2022 · 被引用 25 次
- Discovering Multiple Solutions from a Single Task in Offline Reinforcement LearningTakayuki Osa, Tatsuya HaradaICML 2024 · 被引用 3 次
- Open-Ended Diverse Solution Discovery with Regulated Behavior Patterns for Cross-Domain AdaptationKang Xu, Yan Ma, Bingsheng Wei, Wei LiAAAI 2023 · 被引用 3 次
- Learning a subspace of policies for online adaptation in Reinforcement LearningJean-Baptiste Gaya, Laure Soulier, Ludovic DenoyerICLR 2022 · 被引用 17 次
- Discovering Creative Behaviors through DUPLEX: Diverse Universal Features for Policy ExplorationBorja G. León, Francesco Riccio, Kaushik Subramanian, Peter R. Wurman 等NeurIPS 2024 · 被引用 5 次
