Cold-Start Personalization via Bayesian Adaptive Questioning
Avinandan Bose, Stella Li, Faeze Brahman, Pang Wei Koh, Simon Du, Yulia Tsvetkov, Maryam Fazel, Lin Xiao, Asli Celikyilmaz
Abstract
Cold-start personalization requires inferring preferences from minimal interaction when no userspecific historical data is available. The space of possible preferences is vast, yet users care about only a sparse subset and rarely articulate them upfront; combined with limited interaction budgets, this makes preference elicitation challenging.
Our key insight is that preferences exhibit predictable structure across populations; e.g., users who want detailed explanations often also value worked examples. We propose CAPE (Cold-start Adaptive Preference Elicitation), which learns a structured world model of preference correlations offline from complete profiles, then performs training-free Bayesian inference online to select informative questions and predict complete preference profiles, including dimensions never asked about. Even simple belief model instantiations (e.g., linear regression) substantially outperform end-to-end RL. Across medical, mathematical, social, and commonsense reasoning, CAPE achieves 80.8% alignment with ground-truth user preferences versus 68.5% for RL, requires 3-5× fewer interactions, and adapts twice as often. Our contribution is a principled decomposition of cold-start personalization that makes Bayesian preference elicitation practical at scale for LLM systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He et al.ACL 2026 · 346 citations
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference LearningSriyash Poddar, Yanming Wan, Hamish Ivison, Abhishek Gupta et al.NeurIPS 2024 · 188 citations
- ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationJianghao Lin, Rong Shan, Chenxu Zhu, Kounianhua Du et al.WWW 2024 · 151 citations
Related papers
- PrefDisco: Benchmarking Proactive Personalized ReasoningShuyue Stella Li, Avinandan Bose, Faeze Brahman, Simon S. Du et al.ICLR 2026 · 8 citations
- Supporting High-Stakes Decision Making Through Interactive Preference Elicitation in the Latent SpaceMichael Eichelbeck, Tim Voigt, Matthias AlthoffICLR 2026
- Adaptive Querying with AI Persona PriorsKaizheng Wang, Yuhang Wu, Assaf ZeeviICML 2026
- RPM: Reasoning-Level Personalization for Black-Box Large Language ModelsJieyong Kim, Tongyoung Kim, Soojin Yoon, Jaehyung Kim et al.ICLR 2026 · 3 citations
- Alleviating Cold-start Problem in CTR Prediction with A Variational Embedding Learning FrameworkXiaoxiao Xu, Chen Yang, Qian Yu, Zhiwei Fang et al.WWW 2022 · 41 citations
