Policy Aggregation
Parand A. Alamdari, Soroush Ebadian, Ariel D. Procaccia
摘要
We consider the challenge of AI value alignment with multiple individuals that have different reward functions and optimal policies in an underlying Markov decision process. We formalize this problem as one of policy aggregation, where the goal is to identify a desirable collective policy. We argue that an approach informed by social choice theory is especially suitable. Our key insight is that social choice methods can be reinterpreted by identifying ordinal preferences with volumes of subsets of the state-action occupancy polytope. Building on this insight, we demonstrate that a variety of methods--including approval voting, Borda count, the proportional veto core, and quantile fairness--can be practically applied to policy aggregation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Minimaximalist Approach to Reinforcement Learning from Human FeedbackGokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu 等ICML 2024 · 被引用 147 次
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 被引用 96 次
- Fair Federated Learning via the Proportional Veto CoreBhaskar Ray Chaudhury, Aniket Murhekar, Zhuowen Yuan, Bo Li 等ICML 2024 · 被引用 14 次
- Computing the Proportional Veto CoreEgor Ianovski, Aleksei Y. KondratevAAAI 2021 · 被引用 9 次
相关 Paper
- Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement LearningCheol Woo Kim, Jai Moondra, Shresth Verma, Madeleine Pollack 等ICML 2025
- The Axiomatic Value of Regularization in AI Alignment from Human PreferencesEzgi KorkmazICML 2026
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 被引用 5 次
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 被引用 13 次
