Aligning LLMs by Predicting Preferences from User Writing Samples
Stéphane Aroca-Ouellette, Natalie Mackraz, Barry-John Theobald, Katherine Metcalf
Abstract
Accommodating human preferences is essential for creating aligned LLM agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs acting as writing agents to infer a description of user preferences. Agent alignment then comes from conditioning on the inferred preference description. However, existing methods often produce generic preference descriptions that fail to capture the unique and individualized nature of human preferences. This paper introduces PROSE, a method designed to enhance the precision of preference descriptions inferred from user writing samples. PROSE incorporates two key elements: (1) iterative refinement of inferred preferences, and (2) verification of inferred preferences across multiple user writing samples. We evaluate PROSE with several LLMs (i.e., Qwen2.5 7B and 72B Instruct, GPT-mini, and GPT-4o) on a summarization and an email writing task. We find that PROSE more accurately infers nuanced human preferences, improving the quality of the writing agent's generations over CI-PHER (a state-of-the-art method for inferring preferences) by 33%. Lastly, we demonstrate that ICL and PROSE are complementary methods, and combining them provides up to a 9% improvement over ICL alone. Code: https: //github.com/apple/ml-predict .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
Related papers
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro et al.NeurIPS 2024 · 102 citations
- Learning to summarize user information for personalized reinforcement learning from human feedbackHyunJi Nam, Yanming Wan, Mickel Liu, Peter F. Ahnn et al.ICLR 2026 · 10 citations
- Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized AlignmentWeixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo et al.NeurIPS 2025 · 34 citations
- RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMsJiaxing Wu, Lin Ning, Luyang Liu, Harrison Lee et al.AAAI 2025 · 12 citations
- Black-Box Prompt Optimization: Aligning Large Language Models without Model TrainingJiale Cheng, Xiao Liu, Kehan Zheng, Pei Ke et al.ACL 2024 · 17 citations
