Lune

ICML2026Top-tier venue

Emergent Alignment via Competition

Natalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu, Mirah Shi

2026Year
2Top-tier citations

Abstract

Aligning AI systems with human values remains a fundamental challenge-but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with multiple differently misaligned AI agents. Our key insight is that when the user's utility function lies approximately within the convex hull of the AI agents' utility functions-a condition that becomes weaker as more diverse models become available-strategic competition among the agents can yield outcomes comparable to interacting with a perfectly aligned model.

We model this as a multi-leader Stackelberg game extending Bayesian persuasion to multi-round conversations between differently informed parties. We prove three main results: (1) When perfect alignment would allow the user to learn their Bayes-optimal action, she is also able to obtain her Bayes optimal utility in all equilibria under our convex hull condition; (2) Under a weaker assumption, a nonstrategic user employing quantal response achieves near-optimal utility in all equilibria; (3) When the user selects the best single AI to interact with after an evaluation period, in equilibrium near-optimal utility is guaranteed without any additional distributional assumptions.

We complement the theory with two forms of empirical evidence: First, we test our alignment condition on both synthetic and real-world data. We show that synthetically generated LLM utility functions (produced via perturbations of the same prompt to evaluate instances on a movie recommendation (MovieLens) and ethical judgement (ETHICS) dataset) quickly produce a convex hull that contains a good approximation of a given utility function even when none of the individual LLM utility functions is well aligned. We show similar findings using human and LLM responses on real-world polling data (OpinionQA): a convex hull of LLM opinions can approximate human opinions more accurately than any individual LLM across a wide range of survey questions. Second, we perform simulations of the best-AI selection game using best response dynamics, which show that competition among individually misaligned agents reliably improves user utility when the approximate convex hull assumption is satisfied.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers2

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines