Lune

ICML2026顶会

Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling

Yang Cai, Weiqiang Zheng

2026年份

摘要

Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each prompt, the model produces k≥1k\ge 1 candidate responses and a user selects their preferred one. We introduce (k,f(k))(k,f(k))-robust alignment, which requires the kk-output model to have win rate f(k)f(k) against any other single-output model, and asymptotic universal alignment (U-alignment), which requires f(k)→1f(k)\to 1 as k→∞k\to\infty. Our main result characterizes the optimal convergence rate: there exists a family of single-output policies whose kk-sample product policies achieve U-alignment at rate f(k)=kk+1f(k)=\frac{k}{k+1}, and no method can achieve a faster rate in general. We show that popular post-training methods, including Nash learning from human feedback (NLHF), can fundamentally underutilize the benefits of test-time scaling. Even though NLHF is optimal for k=1k=1, sampling from the resulting (often deterministic) policy cannot guarantee win rates above 12\tfrac{1}{2} except for an arbitrarily small slack. This stems from a lack of output diversity: existing alignment methods can collapse to a single majority-preferred response, making additional samples redundant. In contrast, our approach preserves output diversity and achieves the optimal test-time scaling rate. In particular, we propose a family of symmetric multi-player alignment games and prove that any symmetric Nash equilibrium policy of the (k+1)(k+1)-player alignment game achieves the optimal (k,kk+1)(k,\frac{k}{k+1})-robust alignment. Finally, we provide theoretical convergence guarantees for self-play learning dynamics in these games and extend the framework to opponents that also generate multiple responses.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4776bb7d-fa39-4af1-94e2-a5ce7965bdb3

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖