Personalized LLM Decoding via Contrasting Personal Preference
Hyungjune Bu, ChanJoo Jung, Minjae Kang, Jaehyung Kim
摘要
As large language models (LLMs) are progressively deployed in various real-world applications, personalization of LLMs has become increasingly important. While various approaches to LLM personalization such as prompt-based and training-based methods have been actively explored, the development of effective decoding-time algorithms remains largely overlooked, despite their demonstrated potential. In this paper, we propose COPE (Contrasting Personal Preference), a novel decoding-time approach applied after performing parameter-efficient fine-tuning (PEFT) on user-specific data. Our core idea is to leverage reward-guided decoding specifically for personalization by maximizing each user's implicit reward signal. We evaluate COPE across five open-ended personalized text generation tasks. Our empirical results demonstrate that COPE achieves strong performance, improving personalization by an average of 10.57% in ROUGE-L,without relying on external reward models or additional training procedures. 1 * Equal contribution (listed in alphabetical order). 1 Code is available at https://github.com/ cleverscent/CoPe .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAChanJoo Jung, Jaehyung KimICLR 2026 · 被引用 2 次
- Personalizing LLMs with Binary Feedback: A Preference-Calibrated Optimization FrameworkXilai Ma, Liye Zhao, Weijun Yao, Haibing Di 等ACL 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- T-POP: Test-Time Personalization with Online Preference FeedbackZikun Qu, Min Zhang, Mingze Kong, Xiang Li 等ICML 2026 · 被引用 4 次
- Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningZhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu 等EMNLP 2024 · 被引用 17 次
- Instant Personalized Large Language Model Adaptation via HypernetworkZhaoxuan Tan, Zixuan Zhang, Haoyang Wen, Zheng Li 等ACL 2026 · 被引用 7 次
- CoPL: Collaborative Preference Learning for Personalizing LLMsYoungbin Choi, Seunghyuk Cho, Minjong Lee, MoonJeong Park 等EMNLP 2025
- PAD: Personalized Alignment of LLMs at Decoding-timeRuizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai 等ICLR 2025
