Emergent Alignment via Competition
Natalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu, Mirah Shi
Abstract
Aligning AI systems with human values remains a fundamental challenge-but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with multiple differently misaligned AI agents. Our key insight is that when the user's utility function lies approximately within the convex hull of the AI agents' utility functions-a condition that becomes weaker as more diverse models become available-strategic competition among the agents can yield outcomes comparable to interacting with a perfectly aligned model.
We model this as a multi-leader Stackelberg game extending Bayesian persuasion to multi-round conversations between differently informed parties. We prove three main results: (1) When perfect alignment would allow the user to learn their Bayes-optimal action, she is also able to obtain her Bayes optimal utility in all equilibria under our convex hull condition; (2) Under a weaker assumption, a nonstrategic user employing quantal response achieves near-optimal utility in all equilibria; (3) When the user selects the best single AI to interact with after an evaluation period, in equilibrium near-optimal utility is guaranteed without any additional distributional assumptions.
We complement the theory with two forms of empirical evidence: First, we test our alignment condition on both synthetic and real-world data. We show that synthetically generated LLM utility functions (produced via perturbations of the same prompt to evaluate instances on a movie recommendation (MovieLens) and ethical judgement (ETHICS) dataset) quickly produce a convex hull that contains a good approximation of a given utility function even when none of the individual LLM utility functions is well aligned. We show similar findings using human and LLM responses on real-world polling data (OpinionQA): a convex hull of LLM opinions can approximate human opinions more accurately than any individual LLM across a wide range of survey questions. Second, we perform simulations of the best-AI selection game using best response dynamics, which show that competition among individually misaligned agents reliably improves user utility when the approximate convex hull assumption is satisfied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards Strategic Persuasion with Language ModelsZirui Cheng, Jiaxuan YouICLR 2026 · 11 citations
- The Price of Competitive Information DisclosureSiddhartha Banerjee, Kamesh Munagala, Yiheng Shen, Kangning WangSTOC 2026 · 1 citation
Builds on7
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Scalable AI Safety via Doubly-Efficient DebateJonah Brown-Cohen, Geoffrey Irving, Georgios PiliourasICML 2024 · 42 citations
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri et al.NeurIPS 2025 · 26 citations
- Multi-Sender Persuasion: A Computational PerspectiveSafwan Hossain, Tonghan Wang, Tao Lin, Yiling Chen et al.ICML 2024 · 14 citations
Related papers
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li et al.ICML 2026 · 3 citations
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural NetworksSu Hyeong Lee, Risi Kondor, Richard NgoICML 2026
- Understanding Compliance and Conversion Dynamics in Multi-Agent CollectivesSoohwan Lee, Kyungho LeeCHI 2026 · 1 citation
- The Burden of Interactive Alignment with Inconsistent PreferencesAli ShiraliNeurIPS 2025 · 2 citations
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 11 citations
