DEEPER Insight into Your User: Directed Persona Refinement for Dynamic Persona Modeling
Aili Chen, Chengyu Du, Jiangjie Chen, Jinghan Xu, Yikai Zhang, Siyu Yuan, Zulong Chen, Liangyue Li, Yanghua Xiao
Abstract
To advance personalized applications such as recommendation systems and user behavior prediction, recent research increasingly adopts large language models (LLMs) for humanreadable persona modeling. In dynamic realworld scenarios, effective persona modeling necessitates leveraging streaming behavior data to continually optimize user personas. However, existing methods-whether regenerating personas or incrementally extending them with new behaviors-often fail to achieve sustained improvements in persona quality or future behavior prediction accuracy. To address this, we propose DEEPER, a novel approach for dynamic persona modeling that enables continual persona optimization. Specifically, we enhance the model's direction-search capability through an iterative offline reinforcement learning framework, allowing it to automatically identify effective update directions and optimize personas using discrepancies between user behaviors and model predictions. Extensive experiments on dynamic persona modeling involving 4,800 users across 10 domains highlight DEEPER 's superior persona optimization capabilities, delivering an impressive 32.2% average reduction in user behavior prediction error over four update rounds-outperforming the best baseline by a remarkable 22.92%. 1 Recent studies increasingly utilize Large Language Models (LLMs) (OpenAI, 2023a; Anthropic, 2024; AI@Meta, 2024) for human-readable and interpretable persona modeling, advancing personalized applications like recommendation and behavior prediction. However, most research focuses on generating personas from static historical data, which fail
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b15fa53-a25d-4b30-9a37-83db335f8ad7Cited by top-tier papers2
- Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized AlignmentWeixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo et al.NeurIPS 2025 · 34 citations
- Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue SystemsJihao Zhao, Ding Chen, Zhaoxin Fan, Kerun Xu et al.ACL 2026
Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n SamplingLin Gui, Cristina Garbacea, Victor VeitchNeurIPS 2024 · 138 citations
- Incremental Mobile User Profiling: Reinforcement Learning with Spatial Knowledge Graph for Modeling Event StreamsPengyang Wang, Kunpeng Liu, Lu Jiang, Xiaolin Li et al.KDD 2020 · 72 citations
- UCTopic: Unsupervised Contrastive Learning for Phrase Representations and Topic MiningJiacheng Li, Jingbo Shang, Julian J. McAuleyACL 2022 · 68 citations
- Enhancing Personalized Dialogue Generation with Contrastive Latent Variables: Combining Sparse and Dense PersonaYihong Tang, Bo Wang, Miao Fang, Dongming Zhao et al.ACL 2023 · 9 citations
Related papers
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff et al.NeurIPS 2025 · 51 citations
- Orion: Steering Personalized Web Agents via Global-Micro Profiling and Adaptive Intent TrackingDie Hu, Jingguo Ge, Weitao Tang, He Kong et al.AAAI 2026
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen et al.WWW 2026
- Adaptive Preference Arithmetic: A Personalized Agent with Adaptive Preference Arithmetic for Dynamic Preference ModelingHongyi Nie, Yaqing Wang, Mingyang Zhou, Feiyang Pan et al.NeurIPS 2025 · 1 citation
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath et al.ICML 2026 · 33 citations
