Instance-based Generalization in Reinforcement Learning
Martín Bertrán, Natalia Martínez, Mariano Phielipp, Guillermo Sapiro
摘要
Agents trained via deep reinforcement learning (RL) routinely fail to generalize to unseen environments, even when these share the same underlying dynamics as the training levels. Understanding the generalization properties of RL is one of the challenges of modern machine learning. Towards this goal, we analyze policy learning in the context of Partially Observable Markov Decision Processes (POMDPs) and formalize the dynamics of training levels as instances. We prove that, independently of the exploration strategy, reusing instances introduces significant changes on the effective Markov dynamics the agent observes during training. Maximizing expected rewards impacts the learned belief state of the agent by inducing undesired instance-specific speed-running policies instead of generalizable ones, which are sub-optimal on the training set. We provide generalization bounds to the value gap in train and test environments based on the number of training instances, and use insights based on these to improve performance on unseen levels. We propose training a shared belief representation over an ensemble of specialized policies, from which we compute a consensus policy that is used for data collection, disallowing instance-specific exploitation. We experimentally validate our theory, observations, and the proposed computational solution over the CoinRun benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 被引用 48 次
- When Is Generalizable Reinforcement Learning Tractable?Dhruv Malik, Yuanzhi Li, Pradeep RavikumarNeurIPS 2021 · 被引用 32 次
- Explore to Generalize in Zero-Shot RLEv Zisselman, Itai Lavie, Daniel Soudry, Aviv TamarNeurIPS 2023 · 被引用 26 次
- The Generalization Gap in Offline Reinforcement LearningIshita Mediratta, Qingfei You, Minqi Jiang, Roberta RaileanuICLR 2024 · 被引用 24 次
它引用的顶会 Paper3
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement LearningKimin Lee, Kibok Lee, Jinwoo Shin, Honglak LeeICLR 2020 · 被引用 191 次
- Observational Overfitting in Reinforcement LearningXingyou Song, Yiding Jiang, Stephen Tu, Yilun Du 等ICLR 2020 · 被引用 148 次
相关 Paper
- Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial ObservabilityDibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang 等NeurIPS 2021 · 被引用 176 次
- The Curse of Diversity in Ensemble-Based ExplorationZhixuan Lin, Pierluca D'Oro, Evgenii Nikishin, Aaron C. CourvilleICLR 2024 · 被引用 9 次
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte 等ICML 2023 · 被引用 21 次
- Learning a subspace of policies for online adaptation in Reinforcement LearningJean-Baptiste Gaya, Laure Soulier, Ludovic DenoyerICLR 2022 · 被引用 17 次
- Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationYuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong 等AAAI 2024 · 被引用 4 次
