Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
Frédéric Berdoz, Roger Wattenhofer
Abstract
While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications such as social decision processes. More importantly, existing alignment methods provide no formal guarantees on the safety of such models. Drawing from utility and social choice theory, we provide a novel quantitative definition of alignment in the context of social decision-making. Building on this definition, we introduce probably approximately aligned (i.e., near-optimal) policies, and we derive a sufficient condition for their existence. Lastly, recognizing the practical difficulty of satisfying this condition, we introduce the relaxed concept of safe (i.e., nondestructive) policies, and we propose a simple yet robust method to safeguard the black-box policy of any autonomous agent, ensuring all its actions are verifiably safe for the society.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e4fbf1f-e568-483c-b514-0c2aa7fd4899Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- An Axiomatic Theory of Provably-Fair Welfare-Centric Machine LearningCyrus CousinsNeurIPS 2021 · 39 citations
Related papers
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 11 citations
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 4 citations
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca et al.ICLR 2025
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 14 citations
