Learning Human Objectives by Evaluating Hypothetical Behavior
Siddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg, Jan Leike
Abstract
We seek to align agent behavior with a user's objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states. The user knows the rewards and unsafe states, but querying the user is expensive. To address this challenge, we propose an algorithm that safely and interactively learns a model of the user's reward function. We start with a generative model of initial states and a forward dynamics model trained on off-policy data. Our method uses these models to synthesize hypothetical behaviors, asks the user to label the behaviors with rewards, and trains a neural network to predict the rewards. The key idea is to actively synthesize the hypothetical behaviors from scratch by maximizing tractable proxies for the value of information, without interacting with the environment. We call this method reward query synthesis via trajectory optimization (ReQueST). We evaluate ReQueST with simulated users on a state-based 2D navigation task and the image-based Car Racing video game. The results show that ReQueST significantly outperforms prior methods in learning reward models that transfer to new environments with different initial state distributions. Moreover, ReQueST safely trains the reward model to detect unsafe states, and corrects reward hacking before deploying the agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b0c634d-7343-458f-b709-00d765f46c4bCited by top-tier papers10
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li et al.ICML 2023 · 200 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- Reinforcement Learning Under Moral UncertaintyAdrien Ecoffet, Joel LehmanICML 2021 · 40 citations
- Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance LearningJoseph Early, Tom Bewley, Christine Evers, Sarvapali D. RamchurnNeurIPS 2022 · 22 citations
- Human-AI Shared Control via Policy DissectionQuanyi Li, Zhenghao Peng, Haibin Wu, Lan Feng et al.NeurIPS 2022 · 16 citations
Builds on1
Related papers
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 21 citations
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li et al.ICML 2026 · 3 citations
- Explicable Policy SearchZe Gong, Yu ZhangNeurIPS 2022 · 5 citations
- Safety through feedback in Constrained RLShashank Reddy Chirra, Pradeep Varakantham, Praveen ParuchuriNeurIPS 2024 · 6 citations
