Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber
Abstract
Sure! To liven up your party, you could… or even hire a bartender to make specialty cocktails… Direct Preference Opt. (DPO) Sure! … Whatever you decide, make sure it's something that everyone can enjoy and stay safe! DPO + Persona Tailoring (Ours) Sure! Here are some ideas to liven up your party tonight: ... 10. Have an icebreaker activity Prompt My school is having a cake drive. Would brownies be okay to take? Chosen Persona The user values simplicity and prefers direct, concise answers without additional details Prompt: My school is having a cake drive… Persona: The user values simplicity and prefers direct… Response: Yes, brownies would be a great… Rejected Persona The user is practical, preferring responses that include logistical considerations Response: Yes, brownies would be a great contribution… Persona Inference ( §2, 3) Persona Tailoring ( §4, 5) Prompt: My school is having a cake… Persona: The user values simplicity… 1) Few-shot Prompting 2) Supervised Fine-Tuning 3) Direct Preference Optimization Typical Preference Dataset Can abductive reasoning reveal why users may prefer responses? Prompt: My school is having… Persona: The user values sim… Chosen Response: Yes, brownies would be a great… Rejected Response: Yes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri et al.NeurIPS 2025 · 26 citations
- P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic ChecklistKwangwook Seo, Dongha LeeACL 2026 · 2 citations
- Principled Content Selection to Generate Diverse and Personalized Multi-Document SummariesVishakh Padmakumar, Zichao Wang, David Arbour, Jennifer HealeyACL 2025 · 1 citation
- Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real UsersNishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman et al.ACL 2026 · 1 citation
- BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic FeedbackHyunseo Kim, Sangam Lee, Kwangwook Seo, Dongha LeeICML 2026
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playingLetian Peng, Jingbo ShangNeurIPS 2024 · 21 citations
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen et al.NeurIPS 2025 · 13 citations
- PrefDisco: Benchmarking Proactive Personalized ReasoningShuyue Stella Li, Avinandan Bose, Faeze Brahman, Simon S. Du et al.ICLR 2026 · 8 citations
- Expectation Preference Optimization: Reliable Preference Estimation for Improving the Reasoning Capability of Large Language ModelsZelin Li, Dawei SongEMNLP 2025
- Active Preference Learning for Large Language ModelsWilliam Muldrew, Peter Hayes, Mingtian Zhang, David BarberICML 2024 · 53 citations
