Train Hard, Fight Easy: Robust Meta Reinforcement Learning
Ido Greenberg, Shie Mannor, Gal Chechik, Eli A. Meirom
摘要
A major challenge of reinforcement learning (RL) in real-world applications is the variation between environments, tasks or clients. Meta-RL (MRL) addresses this issue by learning a meta-policy that adapts to new tasks. Standard MRL methods optimize the average return over tasks, but often suffer from poor results in tasks of high risk or difficulty. This limits system reliability since test tasks are not known in advance. In this work, we define a robust MRL objective with a controlled robustness level. Optimization of analogous robust objectives in RL is known to lead to both biased gradients and data inefficiency. We prove that the gradient bias disappears in our proposed MRL framework. The data inefficiency is addressed via the novel Robust Meta RL algorithm (RoML). RoML is a meta-algorithm that generates a robust version of any given MRL algorithm, by identifying and oversampling harder tasks throughout training. We demonstrate that RoML achieves robust returns on multiple navigation and continuous control benchmarks. We test our algorithms on several domains. Section 6.1 considers a navigation problem, where both CVaR-ML and RoML obtain better CVaR returns than their risk-neutral baseline. Furthermore, they learn substantially different navigation policies. Section 6.2 considers several continuous control environments with varying tasks. These environments are challenging for CVaR-ML, which entirely fails to learn. Yet, RoML preserves its effectiveness and consistently improves the robustness of the returns. In addition, Section 6.3 demonstrates that under certain conditions, RoML can be applied to supervised settings as well -providing robust supervised meta-learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Test Time Scaling for Neural ProcessesHyungi Lee, Moonseok Choi, Hyunsu Kim, Kyunghyun Cho 等NeurIPS 2025 · 被引用 1 次
- Probabilistic Performance Guarantees for Multi-Task Reinforcement LearningYannik Schnitzer, Mathias Jackermeier, Alessandro Abate, David ParkerICML 2026 · 被引用 1 次
- Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution TasksJeongmo Kim, Yisak Park, Minung Kim, Seungyul HanICML 2025
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen 等ICML 2026
- Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized EnvironmentsYun Qu, Cheems Wang, Yixiu Mao, Yiqin Lv 等ICML 2025
它引用的顶会 Paper15
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Replay-Guided Adversarial Environment DesignMinqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster 等NeurIPS 2021 · 被引用 148 次
- Adversarially Robust Few-Shot Learning: A Meta-Learning ApproachMicah Goldblum, Liam Fowl, Tom GoldsteinNeurIPS 2020 · 被引用 107 次
- Self-Paced Deep Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenNeurIPS 2020 · 被引用 83 次
相关 Paper
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 被引用 58 次
- Distributionally Adaptive Meta Reinforcement LearningAnurag Ajay, Abhishek Gupta, Dibya Ghosh, Sergey Levine 等NeurIPS 2022 · 被引用 21 次
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Meta-Reinforcement Learning Based on Self-Supervised Task Representation LearningMingyang Wang, Zhenshan Bing, Xiangtong Yao, Shuai Wang 等AAAI 2023 · 被引用 22 次
- Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward EnvironmentsDesik Rengarajan, Sapana Chaudhary, Jaewon Kim, Dileep Kalathil 等NeurIPS 2022 · 被引用 2 次
