Personalized LLM Decoding via Contrasting Personal Preference
Hyungjune Bu, ChanJoo Jung, Minjae Kang, Jaehyung Kim
Abstract
As large language models (LLMs) are progressively deployed in various real-world applications, personalization of LLMs has become increasingly important. While various approaches to LLM personalization such as prompt-based and training-based methods have been actively explored, the development of effective decoding-time algorithms remains largely overlooked, despite their demonstrated potential. In this paper, we propose COPE (Contrasting Personal Preference), a novel decoding-time approach applied after performing parameter-efficient fine-tuning (PEFT) on user-specific data. Our core idea is to leverage reward-guided decoding specifically for personalization by maximizing each user's implicit reward signal. We evaluate COPE across five open-ended personalized text generation tasks. Our empirical results demonstrate that COPE achieves strong performance, improving personalization by an average of 10.57% in ROUGE-L,without relying on external reward models or additional training procedures. 1 * Equal contribution (listed in alphabetical order). 1 Code is available at https://github.com/ cleverscent/CoPe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAChanJoo Jung, Jaehyung KimICLR 2026 · 2 citations
- Personalizing LLMs with Binary Feedback: A Preference-Calibrated Optimization FrameworkXilai Ma, Liye Zhao, Weijun Yao, Haibing Di et al.ACL 2026
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- T-POP: Test-Time Personalization with Online Preference FeedbackZikun Qu, Min Zhang, Mingze Kong, Xiang Li et al.ICML 2026 · 4 citations
- Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningZhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu et al.EMNLP 2024 · 17 citations
- Instant Personalized Large Language Model Adaptation via HypernetworkZhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li et al.ACL 2026 · 7 citations
- CoPL: Collaborative Preference Learning for Personalizing LLMsYoungbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park et al.EMNLP 2025
- PAD: Personalized Alignment of LLMs at Decoding-timeRuizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai et al.ICLR 2025
