Constrained GPI for Zero-Shot Transfer in Reinforcement Learning
Jaekyeom Kim, Seohong Park, Gunhee Kim
摘要
For zero-shot transfer in reinforcement learning where the reward function varies between different tasks, the successor features framework has been one of the popular approaches. However, in this framework, the transfer to new target tasks with generalized policy improvement (GPI) relies on only the source successor features [6] or additional successor features obtained from the function approximators' generalization to novel inputs [12] . The goal of this work is to improve the transfer by more tightly bounding the value approximation errors of successor features on the new target tasks. Given a set of source tasks with their successor features, we present lower and upper bounds on the optimal values for novel task vectors that are expressible as linear combinations of source task vectors. Based on the bounds, we propose constrained GPI as a simple test-time approach that can improve transfer by constraining action-value approximation errors on new target tasks. Through experiments in the Scavenger and Reacher environment with state observations as well as the DeepMind Lab environment with visual observations, we show that the proposed constrained GPI significantly outperforms the prior GPI's transfer performance. Our code and additional information are available at https://jaekyeom.github.io/projects/cgpi/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement LearningYihang Yao, Zuxin Liu, Zhepeng Cen, Jiacheng Zhu 等NeurIPS 2023 · 被引用 24 次
- Intention-Conditioned Flow Occupancy ModelsChongyi Zheng, Seohong Park, Sergey Levine, Benjamin EysenbachICLR 2026 · 被引用 9 次
- Multi-Step Generalized Policy Improvement by Leveraging Approximate ModelsLucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da SilvaNeurIPS 2023 · 被引用 7 次
- Constructing an Optimal Behavior Basis for the Option KeyboardLucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da SilvaNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper5
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
- A New Representation of Successor Features for Transfer across Dissimilar EnvironmentsMajid Abdolshah, Hung Le, Thommen George Karimpanal, Sunil Gupta 等ICML 2021 · 被引用 21 次
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
- Constructing a Good Behavior Basis for Transfer using Generalized Policy UpdatesSafa Alver, Doina PrecupICLR 2022 · 被引用 19 次
相关 Paper
- SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement LearningShuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan 等ICML 2024 · 被引用 7 次
- Distributional Successor Features Enable Zero-Shot Policy OptimizationChuning Zhu, Xinqi Wang, Tyler Han, Simon S. Du 等NeurIPS 2024 · 被引用 11 次
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee 等ICLR 2023 · 被引用 2 次
- Risk-Aware Transfer in Reinforcement Learning using Successor FeaturesMichael Gimelfarb, André Barreto, Scott Sanner, Chi-Guhn LeeNeurIPS 2021 · 被引用 25 次
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney 等ICML 2022 · 被引用 11 次
