Lune

ICML2026顶会

Emergent Alignment via Competition

Natalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu, Mirah Shi

2026年份
2顶会引用

摘要

Aligning AI systems with human values remains a fundamental challenge-but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with multiple differently misaligned AI agents. Our key insight is that when the user's utility function lies approximately within the convex hull of the AI agents' utility functions-a condition that becomes weaker as more diverse models become available-strategic competition among the agents can yield outcomes comparable to interacting with a perfectly aligned model.

We model this as a multi-leader Stackelberg game extending Bayesian persuasion to multi-round conversations between differently informed parties. We prove three main results: (1) When perfect alignment would allow the user to learn their Bayes-optimal action, she is also able to obtain her Bayes optimal utility in all equilibria under our convex hull condition; (2) Under a weaker assumption, a nonstrategic user employing quantal response achieves near-optimal utility in all equilibria; (3) When the user selects the best single AI to interact with after an evaluation period, in equilibrium near-optimal utility is guaranteed without any additional distributional assumptions.

We complement the theory with two forms of empirical evidence: First, we test our alignment condition on both synthetic and real-world data. We show that synthetically generated LLM utility functions (produced via perturbations of the same prompt to evaluate instances on a movie recommendation (MovieLens) and ethical judgement (ETHICS) dataset) quickly produce a convex hull that contains a good approximation of a given utility function even when none of the individual LLM utility functions is well aligned. We show similar findings using human and LLM responses on real-world polling data (OpinionQA): a convex hull of LLM opinions can approximate human opinions more accurately than any individual LLM across a wide range of survey questions. Second, we perform simulations of the best-AI selection game using best response dynamics, which show that competition among individually misaligned agents reliably improves user utility when the approximate convex hull assumption is satisfied.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖