Lune

NeurIPS2025顶会

Inference-Time Personalized Alignment with a Few User Preference Queries

Victor-Alexandru Padurean, Parameswaran Kamalaruban, Nachiket Kotalwar, Alkis Gotovos, Adish Singla

2025年份
5被引次数

摘要

We study the problem of aligning a generative model's response with a user's preferences. Recent works have proposed several different formulations for personalized alignment; however, they either require a large amount of user preference queries or require that the preference be explicitly specified as a text input. In this paper, we propose a novel inference-time personalized alignment method, USERALIGN, that elicits the user's preferences with a few queries as pairwise response comparisons. In particular, USERALIGN builds on the theoretical framework of best-arm identification in logistic bandits and selects a personalized response from a fixed pool of the model's generated responses. The key idea is to consider the user's feedback consistent and noise-free, and incorporate it into the theoretical framework to identify the best response quickly. Experimental results across several tasks, involving personalized text and image generation, showcase the effectiveness of USERALIGN in achieving personalized alignment.

How about a juicy ribeye steak with a side of garlic butter mushrooms for a rich, savory meal? How about a spicy chickpea and vegetable curry with a medley of bell peppers, zucchini, and spinach, served over basmati rice? How about a comforting vegetable risotto with fresh herbs like basil and parsley, where the creamy arborio rice perfectly complements the flavors of sautéed mushrooms, peas, and asparagus?

4 Output How about a comforting vegetable risotto with fresh herbs like basil and parsley, where the creamy arborio rice perfectly complements the flavors of sautéed mushrooms, peas, and asparagus?

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖