Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
Weixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo, Haixiao Liu, Biye Li, Yanyan Zhao, Bing Qin, Ting Liu
Abstract
Personalized alignment is essential for enabling large language models (LLMs) to engage effectively in user-centric dialogue. While recent prompt-based and offline optimization methods offer preliminary solutions, they fall short in coldstart scenarios and long-term personalization due to their inherently static and shallow designs. In this work, we introduce the Reinforcement Learning for Personalized Alignment (RLPA) framework, in which an LLM interacts with a simulated user model to iteratively infer and refine user profiles through dialogue. The training process is guided by a dual-level reward structure: the Profile Reward encourages accurate construction of user representations, while the Response Reward incentivizes generation of responses consistent with the inferred profile. We instantiate RLPA by fine-tuning Qwen-2.5-3B-Instruct, resulting in Qwen-RLPA, which achieves state-of-the-art performance in personalized dialogue. Empirical evaluations demonstrate that Qwen-RLPA consistently outperforms prompting and offline fine-tuning baselines, and even surpasses advanced commercial models such as Claude-3.5 and GPT-4o. Further analysis highlights Qwen-RLPA's robustness in reconciling conflicting user preferences, sustaining long-term personalization and delivering more efficient inference compared to recent reasoning-focused LLMs. These results emphasize the potential of dynamic profile inference as a more effective paradigm for building personalized dialogue systems. Our code is available at: https://github.com/XingYuSSS/RLPA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 273c4a3a-189b-4d53-92ba-81dce1e63a65Cited by top-tier papers6
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue AgentsJiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long et al.ACL 2026 · 5 citations
- P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic ChecklistKwangwook Seo, Dongha LeeACL 2026 · 2 citations
- Latent Inter-User Difference Modeling for LLM PersonalizationYilun Qiu, Tianhao Shi, Xiaoyan Zhao, Fengbin Zhu et al.EMNLP 2025 · 1 citation
- Closing the Expression Gap in LLM Instructions via Socratic QuestioningJianwen Sun, Yukang Feng, Yifan Chang, Chuanhao Li et al.ICML 2026 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- Aligning Large Language Models via Fully Self-Synthetic DataShangjian Yin, Zhepei Wei, Xinyu Zhu, Wei-Lin Chen et al.ACL 2026 · 2 citations
- Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective RewardsHaoxiang Wang, Yong Lin, Wei Xiong, Rui Yang et al.ACL 2024
- PAMDP: Interact to Persona Alignment via a Partially Observable Markov Decision ProcessZhe Yang, Yi Huang, Si Chen, Xiaoting Wu et al.ICLR 2026
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu et al.EMNLP 2025
- "In-Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue LearningChuanqi Cheng, Quan Tu, Wei Wu, Shuo Shang et al.EMNLP 2024 · 3 citations
