Cold-Start Personalization via Bayesian Adaptive Questioning
Avinandan Bose, Stella Li, Faeze Brahman, Pang Wei Koh, Simon Du, Yulia Tsvetkov, Maryam Fazel, Lin Xiao, Asli Celikyilmaz
摘要
Cold-start personalization requires inferring preferences from minimal interaction when no userspecific historical data is available. The space of possible preferences is vast, yet users care about only a sparse subset and rarely articulate them upfront; combined with limited interaction budgets, this makes preference elicitation challenging.
Our key insight is that preferences exhibit predictable structure across populations; e.g., users who want detailed explanations often also value worked examples. We propose CAPE (Cold-start Adaptive Preference Elicitation), which learns a structured world model of preference correlations offline from complete profiles, then performs training-free Bayesian inference online to select informative questions and predict complete preference profiles, including dimensions never asked about. Even simple belief model instantiations (e.g., linear regression) substantially outperform end-to-end RL. Across medical, mathematical, social, and commonsense reasoning, CAPE achieves 80.8% alignment with ground-truth user preferences versus 68.5% for RL, requires 3-5× fewer interactions, and adapts twice as often. Our contribution is a principled decomposition of cold-start personalization that makes Bayesian preference elicitation practical at scale for LLM systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He 等ACL 2026 · 被引用 346 次
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta 等NeurIPS 2024 · 被引用 188 次
- ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationJianghao Lin, Rong Shan, Chenxu Zhu, Kounianhua Du 等WWW 2024 · 被引用 151 次
相关 Paper
- PrefDisco: Benchmarking Proactive Personalized ReasoningShuyue Stella Li, Avinandan Bose, Faeze Brahman, Simon S. Du 等ICLR 2026 · 被引用 8 次
- Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent SpaceMichael Eichelbeck, Tim Voigt, Matthias AlthoffICLR 2026
- Adaptive Querying with AI Persona PriorsKaizheng Wang, Yuhang Wu, Assaf ZeeviICML 2026
- RPM: Reasoning-Level Personalization for Black-Box Large Language ModelsJieyong Kim, Tongyoung Kim, Soojin Yoon, Jaehyung Kim 等ICLR 2026 · 被引用 3 次
- Alleviating Cold-start Problem in CTR Prediction with A Variational Embedding Learning FrameworkXiaoxiao Xu, Chen Yang, Qian Yu, Zhiwei Fang 等WWW 2022 · 被引用 41 次
