Learning Human Objectives by Evaluating Hypothetical Behavior
Siddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg, Jan Leike
摘要
We seek to align agent behavior with a user's objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states. The user knows the rewards and unsafe states, but querying the user is expensive. To address this challenge, we propose an algorithm that safely and interactively learns a model of the user's reward function. We start with a generative model of initial states and a forward dynamics model trained on off-policy data. Our method uses these models to synthesize hypothetical behaviors, asks the user to label the behaviors with rewards, and trains a neural network to predict the rewards. The key idea is to actively synthesize the hypothetical behaviors from scratch by maximizing tractable proxies for the value of information, without interacting with the environment. We call this method reward query synthesis via trajectory optimization (ReQueST). We evaluate ReQueST with simulated users on a state-based 2D navigation task and the image-based Car Racing video game. The results show that ReQueST significantly outperforms prior methods in learning reward models that transfer to new environments with different initial state distributions. Moreover, ReQueST safely trains the reward model to detect unsafe states, and corrects reward hacking before deploying the agent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li 等ICML 2023 · 被引用 200 次
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi 等ICML 2024 · 被引用 145 次
- Reinforcement Learning Under Moral UncertaintyAdrien Ecoffet, Joel LehmanICML 2021 · 被引用 40 次
- Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance LearningJoseph Early, Tom Bewley, Christine Evers, Sarvapali D. RamchurnNeurIPS 2022 · 被引用 22 次
- Human-AI Shared Control via Policy DissectionQuanyi Li, Zhenghao Peng, Haibin Wu, Lan Feng 等NeurIPS 2022 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li 等ICML 2026 · 被引用 3 次
- Explicable Policy SearchZe Gong, Yu ZhangNeurIPS 2022 · 被引用 5 次
- Safety through feedback in Constrained RLShashank Reddy Chirra, Pradeep Varakantham, Praveen ParuchuriNeurIPS 2024 · 被引用 6 次
