Measuring memorization in RLHF for code completion
Jamie Hayes, Ilia Shumailov, William P. Porter, Aneesh Pappu
Abstract
Reinforcement learning with human feedback (RLHF) has become the dominant method to align large models to user preferences. Unlike fine-tuning, for which there are many studies regarding training data memorization, it is not clear how memorization is affected by or introduced in the RLHF alignment process. Understanding this relationship is important as real user data may be collected and used to align large models; if user data is memorized during RLHF and later regurgitated, this could raise privacy concerns. In addition to RLHF, other methods such as Direct Preference Optimization (DPO) and ΨPO have gained popularity for learning directly from human preferences, removing the need for optimizing intermediary reward models with reinforcement learning. In this work, we analyze how training data memorization can surface and propagate through each phase of RLHF and direct preference learning. We focus our study on code completion models, as code completion is one of the most popular use cases for large language models. We find that RLHF significantly decreases the chance that data used for reward modeling and reinforcement learning is memorized in comparison to directly fine-tuning on this data, but that examples already memorized during the fine-tuning stage of RLHF, will, in the majority of cases, remain memorized after RLHF. In contrast, we find that aligning by learning directly from human preference data via a special case of ΨPO, Identity Preference Optimization (IPO), increases the likelihood that training data is regurgitated compared to RLHF. Our work suggests that RLHF, as opposed to direct preference learning, is a safer way to mitigate the risk of regurgitating sensitive preference data when aligning large language models. We find our conclusions are robust across multiple code completion datasets, tasks, and model scales.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d36d6347-ed8c-4eb0-a01b-7795ffef0624Cited by top-tier papers3
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong et al.ICLR 2026 · 150 citations
- Extracting alignment data in open modelsFederico Barbero, Xiangming Gu, Christopher A. Choquette Choo, Chawin Sitawarin et al.ICML 2026 · 9 citations
- Beyond Indistinguishability: Measuring Extraction Risk in LLM APIsRuixuan Liu, David Evans, Li XiongS&P 2026 · 1 citation
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
Related papers
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye et al.ICML 2024 · 274 citations
- Active Preference Learning for Large Language ModelsWilliam Muldrew, Peter Hayes, Mingtian Zhang, David BarberICML 2024 · 53 citations
- Explicit Preference Optimization: No Need for an Implicit Reward ModelXiangkun Hu, Lemin Kong, Tong He, David WipfICML 2025
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelKai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge et al.CVPR 2024 · 34 citations
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu et al.NeurIPS 2024 · 119 citations
