Optimistic Linear Support and Successor Features as a Basis for Optimal Policy Transfer
Lucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da Silva
摘要
In many real-world applications, reinforcement learning (RL) agents might have to solve multiple tasks, each one typically modeled via a reward function. If reward functions are expressed linearly, and the agent has previously learned a set of policies for different tasks, successor features (SFs) can be exploited to combine such policies and identify reasonable solutions for new problems. However, the identified solutions are not guaranteed to be optimal. We introduce a novel algorithm that addresses this limitation. It allows RL agents to combine existing policies and directly identify optimal policies for arbitrary new problems, without requiring any further interac-tions with the environment. We first show (under mild assumptions) that the transfer learning problem tackled by SFs is equivalent to the problem of learning to optimize multiple objectives in RL. We then introduce an SF-based extension of the Optimistic Linear Support algorithm to learn a set of policies whose SFs form a convex coverage set. We prove that policies in this set can be combined via generalized policy improvement to construct optimal behaviors for any new linearly-expressible tasks, without requiring any additional training samples. We empirically show that our method outperforms state-of-the-art competing algorithms both in discrete and continuous domains under value function approximation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Distributional Pareto-Optimal Multi-Objective Reinforcement LearningXin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian 等NeurIPS 2023 · 被引用 46 次
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards 等NeurIPS 2024 · 被引用 14 次
- Task Adaptation from Skills: Information Geometry, Disentanglement, and New Objectives for Unsupervised Reinforcement LearningYucheng Yang, Tianyi Zhou, Qiang He, Lei Han 等ICLR 2024 · 被引用 13 次
- Constrained GPI for Zero-Shot Transfer in Reinforcement LearningJaekyeom Kim, Seohong Park, Gunhee KimNeurIPS 2022 · 被引用 10 次
- On the Occupancy Measure of Non-Markovian Policies in Continuous MDPsRomain Laroche, Remi Tachet des CombesICML 2023 · 被引用 7 次
它引用的顶会 Paper9
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- An Equivalence between Loss Functions and Non-Uniform Sampling in Experience ReplayScott Fujimoto, David Meger, Doina PrecupNeurIPS 2020 · 被引用 85 次
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 被引用 71 次
- The Information Geometry of Unsupervised Reinforcement LearningBenjamin Eysenbach, Ruslan Salakhutdinov, Sergey LevineICLR 2022 · 被引用 41 次
- Discovering a set of policies for the worst case rewardTom Zahavy, André Barreto, Daniel J. Mankowitz, Shaobo Hou 等ICLR 2021 · 被引用 26 次
相关 Paper
- SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement LearningShuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan 等ICML 2024 · 被引用 7 次
- Constructing a Good Behavior Basis for Transfer using Generalized Policy UpdatesSafa Alver, Doina PrecupICLR 2022 · 被引用 19 次
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
- Constructing an Optimal Behavior Basis for the Option KeyboardLucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da SilvaNeurIPS 2025 · 被引用 4 次
- Risk-Aware Transfer in Reinforcement Learning using Successor FeaturesMichael Gimelfarb, André Barreto, Scott Sanner, Chi-Guhn LeeNeurIPS 2021 · 被引用 25 次
