Value Alignment Verification
Daniel S. Brown, Jordan Schneider, Anca D. Dragan, Scott Niekum
Abstract
As humans interact with autonomous agents to perform increasingly complicated, potentially risky tasks, it is important to be able to efficiently evaluate an agent's performance and correctness. In this paper we formalize and theoretically analyze the problem of efficient value alignment verification: how to efficiently test whether the behavior of another agent is aligned with a human's values. The goal is to construct a kind of "driver's test" that a human can give to any agent which will verify value alignment via a minimal number of queries. We study alignment verification problems with both idealized humans that have an explicit reward function as well as problems where they have implicit values. We analyze verification of exact value alignment for rational agents and propose and analyze heuristic and approximate value alignment verification tests in a wide range of gridworlds and a continuous autonomous driving domain. Finally, we prove that there exist sufficient conditions such that we can verify exact and approximate alignment across an infinite set of test environments via a constantquery-complexity alignment test.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6c8e23a-5d63-4d5b-b297-bd96074af901Cited by top-tier papers3
- Identifiability and generalizability from multiple experts in Inverse Reinforcement LearningPaul Rolland, Luca Viano, Norman Schürhoff, Boris Nikolov et al.NeurIPS 2022 · 22 citations
- Curriculum Design for Teaching via Demonstrations: Theory and ApplicationsGaurav Yengera, Rati Devidze, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 10 citations
- Reward Model Evaluation via Automatically-Ranked Policy AlignmentAoran Wang, Lei Ou, Yang Yu, Zongzhang ZhangAAAI 2026
Builds on2
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Safe Imitation Learning via Fast Bayesian Reward Inference from PreferencesDaniel S. Brown, Russell Coleman, Ravi Srinivasan, Scott NiekumICML 2020 · 113 citations
Related papers
- Belief-Driven Value Alignment for Human-Robot CollaborationSaisai Li, Bing Shi, Yiming Xia, Xiao SuAAAI 2026
- Learning Human-like Representations to Enable Learning Human ValuesAndrea Wynn, Ilia Sucholutsky, Tom GriffithsNeurIPS 2024 · 11 citations
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu et al.ICLR 2025
- Observation Interference in Partially Observable Assistance GamesScott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart RussellICML 2025
- The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and AutonomyWilliam Overman, Mohsen BayatiICML 2026 · 7 citations
