Moral Machine or Tyranny of the Majority?
Michael Feffer, Hoda Heidari, Zachary C. Lipton
Abstract
With Artificial Intelligence systems increasingly applied in consequential domains, researchers have begun to ask how these systems ought to act in ethically charged situations where even humans lack consensus. In the Moral Machine project, researchers crowdsourced answers to "Trolley Problems" concerning autonomous vehicles. Subsequently, Noothigattu et al. (2018) proposed inferring linear functions that approximate each individual's preferences and aggregating these linear models by averaging parameters across the population. In this paper, we examine this averaging mechanism, focusing on fairness concerns in the presence of strategic effects. We investigate a simple setting where the population consists of two groups, with the minority constituting an α < 0.5 share of the population. To simplify the analysis, we consider the extreme case in which within-group preferences are homogeneous. Focusing on the fraction of contested cases where the minority group prevails, we make the following observations: (a) even when all parties report their preferences truthfully, the fraction of disputes where the minority prevails is less than proportionate in α; (b) the degree of sub-proportionality grows more severe as the level of disagreement between the groups increases; (c) when parties report preferences strategically, pure strategy equilibria do not always exist; and (d) whenever a pure strategy equilibrium exists, the majority group prevails 100% of the time. These findings raise concerns about stability and fairness of preference vector averaging as a mechanism for aggregating diverging voices. Finally, we discuss alternatives, including randomized dictatorship and median-based mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Proportional Aggregation of Preferences for Sequential Decision MakingNikhil Chandak, Shashwat Goel, Dominik PetersAAAI 2024 · 22 citations
- Wikibench: Community-Driven Data Curation for AI Evaluation on WikipediaTzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim et al.CHI 2024 · 19 citations
- Can AI Model the Complexities of Human Moral Decision-making? A Qualitative Study of Kidney Allocation DecisionsVijay Keswani, Vincent Conitzer, Walter Sinnott-Armstrong, Breanna K. Nguyen et al.CHI 2025 · 11 citations
- Benchmarking Overton Pluralism in LLMsElinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei et al.ICLR 2026 · 9 citations
- Swap-guided Preference Learning for Personalized Reinforcement Learning from Human FeedbackGihoon Kim, Euntai KimICLR 2026 · 5 citations
Builds on1
Related papers
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 5 citations
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 11 citations
- Random Rank: The One and Only Strategyproof and Proportionally Fair Randomized Facility Location MechanismHaris Aziz, Alexander Lam, Mashbat Suzuki, Toby WalshNeurIPS 2022 · 11 citations
- Truthful Aggregation of Budget Proposals with Proportionality GuaranteesIoannis Caragiannis, George Christodoulou, Nicos ProtopapasAAAI 2022 · 21 citations
- Investigating LLM-Powered Dissenting Minority Support in Power-Imbalanced Group Decision-Making: Counterargument and Mediation as Intervention StrategiesSoohwan Lee, Seoyeong Hwang, Mingyu Kim, Dajung Kim et al.CSCW 2026
