Emergent Alignment via Competition
Natalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu, Mirah Shi
摘要
Aligning AI systems with human values remains a fundamental challenge-but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with multiple differently misaligned AI agents. Our key insight is that when the user's utility function lies approximately within the convex hull of the AI agents' utility functions-a condition that becomes weaker as more diverse models become available-strategic competition among the agents can yield outcomes comparable to interacting with a perfectly aligned model.
We model this as a multi-leader Stackelberg game extending Bayesian persuasion to multi-round conversations between differently informed parties. We prove three main results: (1) When perfect alignment would allow the user to learn their Bayes-optimal action, she is also able to obtain her Bayes optimal utility in all equilibria under our convex hull condition; (2) Under a weaker assumption, a nonstrategic user employing quantal response achieves near-optimal utility in all equilibria; (3) When the user selects the best single AI to interact with after an evaluation period, in equilibrium near-optimal utility is guaranteed without any additional distributional assumptions.
We complement the theory with two forms of empirical evidence: First, we test our alignment condition on both synthetic and real-world data. We show that synthetically generated LLM utility functions (produced via perturbations of the same prompt to evaluate instances on a movie recommendation (MovieLens) and ethical judgement (ETHICS) dataset) quickly produce a convex hull that contains a good approximation of a given utility function even when none of the individual LLM utility functions is well aligned. We show similar findings using human and LLM responses on real-world polling data (OpinionQA): a convex hull of LLM opinions can approximate human opinions more accurately than any individual LLM across a wide range of survey questions. Second, we perform simulations of the best-AI selection game using best response dynamics, which show that competition among individually misaligned agents reliably improves user utility when the approximate convex hull assumption is satisfied.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Strategic Persuasion with Language ModelsZirui Cheng, Jiaxuan YouICLR 2026 · 被引用 11 次
- The Price of Competitive Information DisclosureSiddhartha Banerjee, Kamesh Munagala, Yiheng Shen, Kangning WangSTOC 2026 · 被引用 1 次
它引用的顶会 Paper7
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Scalable AI Safety via Doubly-Efficient DebateJonah Brown-Cohen, Geoffrey Irving, Georgios PiliourasICML 2024 · 被引用 42 次
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri 等NeurIPS 2025 · 被引用 26 次
- Multi-Sender Persuasion: A Computational PerspectiveSafwan Hossain, Tonghan Wang, Tao Lin, Yiling Chen 等ICML 2024 · 被引用 14 次
相关 Paper
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li 等ICML 2026 · 被引用 3 次
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural NetworksSu Hyeong Lee, Risi Kondor, Richard NgoICML 2026
- Understanding Compliance and Conversion Dynamics in Multi-Agent CollectivesSoohwan Lee, Kyungho LeeCHI 2026 · 被引用 1 次
- The Burden of Interactive Alignment with Inconsistent PreferencesAli ShiraliNeurIPS 2025 · 被引用 2 次
- Policy AggregationParand A. Alamdari, Soroush Ebadian, Ariel D. ProcacciaNeurIPS 2024 · 被引用 11 次
