Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, Adam Gleave
摘要
It is often very challenging to manually design reward functions for complex, real-world tasks. To solve this, one can instead use reward learning to infer a reward function from data. However, there are often multiple reward functions that fit the data equally well, even in the infinite-data limit. This means that the reward function is only partially identifiable. In this work, we formally characterise the partial identifiability of the reward function given several popular reward learning data sources, including expert demonstrations and trajectory comparisons. We also analyse the impact of this partial identifiability for several downstream tasks, such as policy optimisation. We unify our results in a framework for comparing data sources and downstream tasks by their invariances, with implications for the design and selection of data sources for reward learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand Siththaranjan, Cassidy Laidlaw, Dylan Hadfield-MenellICLR 2024 · 被引用 112 次
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward ModelingYuchun Miao, Sen Zhang, Liang Ding, Rong Bao 等NeurIPS 2024 · 被引用 108 次
- Goodhart's Law in Reinforcement LearningJacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer 等ICLR 2024 · 被引用 22 次
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov 等NeurIPS 2022 · 被引用 22 次
- Identifiability and Generalizability in Constrained Inverse Reinforcement LearningAndreas Schlaginhaufen, Maryam KamgarpourICML 2023 · 被引用 18 次
它引用的顶会 Paper3
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 被引用 219 次
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent 等NeurIPS 2021 · 被引用 106 次
- Identifiability in inverse reinforcement learningHaoyang Cao, Samuel N. Cohen, Lukasz SzpruchNeurIPS 2021 · 被引用 72 次
相关 Paper
- Environment Design for Inverse Reinforcement LearningThomas Kleine Buening, Victor Villin, Christos DimitrakakisICML 2024 · 被引用 5 次
- Reward Identification in Inverse Reinforcement LearningKuno Kim, Shivam Garg, Kirankumar Shiragur, Stefano ErmonICML 2021 · 被引用 43 次
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 被引用 16 次
- Causal Imitation Learning With Unobserved ConfoundersJunzhe Zhang, Daniel Kumor, Elias BareinboimNeurIPS 2020 · 被引用 86 次
- A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and FeedbackKihyun Kim, Jiawei Zhang, Asuman E. Ozdaglar, Pablo A. ParriloICML 2024 · 被引用 2 次
