Value Alignment Verification
Daniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott Niekum
摘要
As humans interact with autonomous agents to perform increasingly complicated, potentially risky tasks, it is important to be able to efficiently evaluate an agent's performance and correctness. In this paper we formalize and theoretically analyze the problem of efficient value alignment verification: how to efficiently test whether the behavior of another agent is aligned with a human's values. The goal is to construct a kind of "driver's test" that a human can give to any agent which will verify value alignment via a minimal number of queries. We study alignment verification problems with both idealized humans that have an explicit reward function as well as problems where they have implicit values. We analyze verification of exact value alignment for rational agents and propose and analyze heuristic and approximate value alignment verification tests in a wide range of gridworlds and a continuous autonomous driving domain. Finally, we prove that there exist sufficient conditions such that we can verify exact and approximate alignment across an infinite set of test environments via a constantquery-complexity alignment test.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov 等NeurIPS 2022 · 被引用 22 次
- Curriculum Design for Teaching via Demonstrations: Theory and ApplicationsGaurav Yengera, Rati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 被引用 10 次
- Reward Model Evaluation via Automatically-Ranked Policy AlignmentAoran Wang, Lei Ou, Yang Yu, Zongzhang ZhangAAAI 2026
它引用的顶会 Paper2
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 被引用 113 次
相关 Paper
- Belief-Driven Value Alignment for Human-Robot CollaborationSaisai Li, Bing Shi, Yiming Xia, Xiao SuAAAI 2026
- Learning Human-like Representations to Enable Learning Human ValuesAndrea Wynn, Ilia Sucholutsky, Tom GriffithsNeurIPS 2024 · 被引用 11 次
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu 等ICLR 2025
- Observation Interference in Partially Observable Assistance GamesScott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart RussellICML 2025
- The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and AutonomyWilliam Overman, Mohsen BayatiICML 2026 · 被引用 7 次
