Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
Frédéric Berdoz, Roger Wattenhofer
摘要
While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications such as social decision processes. More importantly, existing alignment methods provide no formal guarantees on the safety of such models. Drawing from utility and social choice theory, we provide a novel quantitative definition of alignment in the context of social decision-making. Building on this definition, we introduce probably approximately aligned (i.e., near-optimal) policies, and we derive a sufficient condition for their existence. Lastly, recognizing the practical difficulty of satisfying this condition, we introduce the relaxed concept of safe (i.e., nondestructive) policies, and we propose a simple yet robust method to safeguard the black-box policy of any autonomous agent, ensuring all its actions are verifiably safe for the society.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 被引用 120 次
- An Axiomatic Theory of Provably-Fair Welfare-Centric Machine LearningCyrus CousinsNeurIPS 2021 · 被引用 39 次
相关 Paper
- Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained OptimizationTuan Nguyen, Long Tran-ThanhICML 2026
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 被引用 11 次
- Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation MismatchMalek Mechergui, Sarath SreedharanNeurIPS 2024 · 被引用 4 次
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca 等ICLR 2025
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 被引用 14 次
