P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
Kwangwook Seo, Dongha Lee
Abstract
Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with individual preferences. However, existing approaches largely treat user context as a static or implicit conditioning signal, failing to capture the dynamic and multi-faceted nature of human judgment. In this paper, we propose P-CHECK, a novel personalized reward modeling framework, designed to train a plugand-play checklist generator that synthesizes dynamic evaluation criteria for guiding the reward prediction. To better align these checklists with personalized nuances, we introduce Preference-Contrastive Criterion Weighting, a training strategy that assigns saliency scores to criteria based on their discriminative power for personalized judgment. We conduct extensive experiments and demonstrate that P-CHECK not only improves reward accuracy but also enhances downstream personalized generation, and remains robust in OOD scenarios. [CODE]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09c3b5c1-3eab-4df0-93da-967bf0c4b776Builds on25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Dynamic and Generalizable Process Reward ModelingZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng et al.ACL 2025 · 13 citations
- CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward ModelingDengcan Liu, Fengkai Yang, Xiaohan Wang, Shurui Yan et al.KDD 2026 · 12 citations
- SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMsMichael J. Ryan, Omar Shaikh, Aditri Bhagirath, Daniel Frees et al.ACL 2025 · 14 citations
- Check Your Work: Structured Checklist Feedback for Improving Large Language ModelsJonathan Cook, Tim Rocktäschel, Jakob Nicolaus Foerster, Dennis Aumiller et al.ACL 2026
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen et al.ICLR 2026 · 5 citations
