Risk-Aware Transfer in Reinforcement Learning using Successor Features
Michael Gimelfarb, André Barreto, Scott Sanner, Chi-Guhn Lee
摘要
Sample efficiency and risk-awareness are central to the development of practical reinforcement learning (RL) for complex decision-making. The former can be addressed by transfer learning and the latter by optimizing some utility function of the return. However, the problem of transferring skills in a risk-aware manner is not well-understood. In this paper, we address the problem of risk-aware policy transfer between tasks in a common domain that differ only in their reward functions, in which risk is measured by the variance of reward streams. Our approach begins by extending the idea of generalized policy improvement to maximize entropic utilities, thus extending policy improvement via dynamic programming to sets of policies and levels of risk-aversion. Next, we extend the idea of successor features (SF), a value function representation that decouples the environment dynamics from the rewards, to capture the variance of returns. Our resulting risk-aware successor features (RaSF) integrate seamlessly within the RL framework, inherit the superior task generalization ability of SFs, and incorporate risk-awareness into the decision-making. Experiments on a discrete navigation domain and control of a simulated robotic arm demonstrate the ability of RaSFs to outperform alternative methods including SFs, when taking the risk of the learned policies into account.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 被引用 21 次
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang 等ICML 2024 · 被引用 11 次
- Multi-Step Generalized Policy Improvement by Leveraging Approximate ModelsLucas Nunes Alegre, Ana L. C. Bazzan, Ann Nowé, Bruno C. da SilvaNeurIPS 2023 · 被引用 7 次
- Constructing an Optimal Behavior Basis for the Option KeyboardLucas N. Alegre, Ana L. C. Bazzan, André Barreto, Bruno C. da SilvaNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper4
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 等NeurIPS 2020 · 被引用 109 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Mean-Variance Policy Iteration for Risk-Averse Reinforcement LearningShangtong Zhang, Bo Liu, Shimon WhitesonAAAI 2021 · 被引用 44 次
- Variance Penalized On-Policy and Off-Policy Actor-CriticArushi Jain, Gandharv Patil, Ayush Jain, Khimya Khetarpal 等AAAI 2021 · 被引用 11 次
相关 Paper
- Policy Caches with Successor FeaturesMark W. Nemecek, Ron ParrICML 2021 · 被引用 19 次
- SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement LearningShuai Zhang, Heshan Devaka Fernando, Miao Liu, Keerthiram Murugesan 等ICML 2024 · 被引用 7 次
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee 等ICLR 2023 · 被引用 2 次
- A New Representation of Successor Features for Transfer across Dissimilar EnvironmentsMajid Abdolshah, Hung Le, Thommen George Karimpanal, Sunil Gupta 等ICML 2021 · 被引用 21 次
- Bridging Successor Measure and Online Policy Learning with Flow Matching-Based RepresentationsHaosen Shi, Jianda Chen, Sinno Jialin PanICLR 2026
