Privately Aligning Language Models with Reinforcement Learning
Fan Wu, Huseyin A. Inan, Arturs Backurs, Varun Chandrasekaran, Janardhan Kulkarni, Robert Sim
Abstract
Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGPT. In this work, we initiate the study of privacy-preserving alignment of LLMs through Differential Privacy (DP) in conjunction with RL. Following the influential work of Ziegler et al. ( 2020 ), we study two dominant paradigms: (i) alignment via RL without human in the loop (e.g., positive review generation) and (ii) alignment via RL from human feedback (RLHF) (e.g., summarization in a human-preferred way). We give a new DP framework to achieve alignment via RL, and prove its correctness. Our experimental results validate the effectiveness of our approach, offering competitive utility while ensuring strong privacy protections. * This work was carried out as part of an internship at Microsoft Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f457c623-04d3-44eb-a11d-db98de92f34aCited by top-tier papers9
- Privacy-Preserving Instructions for Aligning Large Language ModelsDa Yu, Peter Kairouz, Sewoong Oh, Zheng XuICML 2024 · 41 citations
- On the Sample Complexity of Differentially Private Policy OptimizationYi He, Xingyu ZhouNeurIPS 2025 · 3 citations
- Differentially Private Reinforcement Learning with Self-PlayDan Qiao, Yu-Xiang WangNeurIPS 2024 · 3 citations
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted ControlYuzheng Hu, Ryan McKenna, Da Yu, Shanshan Wu et al.ICML 2026 · 1 citation
- Autoregressive Direct Preference OptimizationMasanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki et al.ICML 2026 · 1 citation
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- Private Direct Preference Optimization for LLM AlignmentYangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao et al.CCS 2026
- Differentially Private Preference Data Synthesis for Large Language Model AlignmentFengyu Gao, Jing YangICML 2026
- Aligning Large Language Models through Synthetic FeedbackSungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang et al.EMNLP 2023 · 13 citations
- LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveBowen Jin, Jinsung Yoon, Zhen Qin, Ziqi Wang et al.ICML 2025
- Measuring memorization in RLHF for code completionJamie Hayes, Ilia Shumailov, William P. Porter, Aneesh PappuICLR 2025
