Constructing a Good Behavior Basis for Transfer using Generalized Policy Updates
Safa Alver, Doina Precup
摘要
We study the problem of learning a good set of policies, so that when combined together, they can solve a wide variety of unseen reinforcement learning tasks with no or very little new data. Specifically, we consider the framework of generalized policy evaluation and improvement, in which the rewards for all tasks of interest are assumed to be expressible as a linear combination of a fixed set of features. We show theoretically that, under certain assumptions, having access to a specific set of diverse policies, which we call a set of independent policies, can allow for instantaneously achieving high-level performance on all possible downstream tasks which are typically more complex than the ones on which the agent was trained. Based on this theoretical analysis, we propose a simple algorithm that iteratively constructs this set of policies. In addition to empirically validating our theoretical results, we compare our approach with recently proposed diverse policy set construction methods and show that, while others fail, our approach is able to build a behavior basis that enables instantaneous transfer to all possible downstream tasks. We also show empirically that having access to a set of independent policies can better bootstrap the learning process on downstream tasks where the new reward function cannot be described as a linear combination of the features. Finally, we demonstrate how this policy set can be useful in a lifelong reinforcement learning setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 被引用 23 次
- Skill Machines: Temporal Logic Skill Composition in Reinforcement LearningGeraud Nangue Tasse, Devon Jarvis, Steven James, Benjamin RosmanICLR 2024 · 被引用 12 次
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney 等ICML 2022 · 被引用 11 次
- Constrained GPI for Zero-Shot Transfer in Reinforcement LearningJaekyeom Kim, Seohong Park, Gunhee KimNeurIPS 2022 · 被引用 10 次
它引用的顶会 Paper5
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 被引用 71 次
- Discovering a set of policies for the worst case rewardTom Zahavy, André Barreto, Daniel J. Mankowitz, Shaobo Hou 等ICLR 2021 · 被引用 26 次
- Generalisation in Lifelong Reinforcement Learning through Logical CompositionGeraud Nangue Tasse, Steven James, Benjamin RosmanICLR 2022 · 被引用 23 次
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
相关 Paper
- Constructing an Optimal Behavior Basis for the Option KeyboardLucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da SilvaNeurIPS 2025 · 被引用 4 次
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 被引用 109 次
- Provably Efficient Lifelong Reinforcement Learning with Linear RepresentationSanae Amani, Lin Yang, Ching-An ChengICLR 2023
- Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter MergingYajat Yadav, Zhiyuan Zhou, Andrew Wagenmaker, Karl Pertsch 等ICLR 2026 · 被引用 10 次
- Toward Robust Long Range Policy TransferWei-Cheng Tseng, Jin-Siang Lin, Yao-Min Feng, Min SunAAAI 2021 · 被引用 8 次
