CPPO: Continual Learning for Reinforcement Learning with Human Feedback
Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, Ruifeng Xu
Abstract
The approach of Reinforcement Learning from Human Feedback (RLHF) is widely used for enhancing pre-trained Language Models (LM), enabling them to better align with human preferences. Existing RLHF-based LMs however require complete retraining whenever new queries or feedback are introduced, as human preferences may differ across different domains or topics. LM retraining is often impracticable in most real-world scenarios, due to the substantial time and computational costs involved, as well as data privacy concerns. To address this limitation, we propose Continual Proximal Policy Optimization (CPPO), a novel method that is able to continually align LM with dynamic human preferences. Specifically, CPPO adopts a weighting strategy to decide which samples should be utilized for enhancing policy learning and which should be used for solidifying past experiences. This seeks a good trade-off between policy learning and knowledge retention. Our experimental results show that CPPO outperforms strong Continuous learning (CL) baselines when it comes to consistently aligning with human preferences. Furthermore, compared to PPO, CPPO offers more efficient and stable learning in non-continual scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7196c1de-9b20-4f1e-b92a-5b0be2635843Cited by top-tier papers9
- ReDit: Reward Dithering for Improved LLM Policy OptimizationChenxing Wei, Jiarui Yu, Ying He, Hande Dong et al.NeurIPS 2025 · 14 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference OptimizationJunsong Li, Jie Zhou, Bihao Zhan, Yutao Yang et al.AAAI 2026 · 3 citations
- GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement LearningHan Zhang, RuibinZheng, ZEXUAN YI, Zhuo Zhang et al.ICLR 2026 · 3 citations
- φ-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal ModelsThanh-Dat Truong, Huu-Thien Tran, Jackson David Cothren, Bhiksha Raj et al.CVPR 2026 · 2 citations
Builds on8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
Related papers
- OPPO: Accelerating PPO-based RLHF via Pipeline OverlapKaizhuo Yan, Yingjie Yu, Yifan Yu, Haizhong Zheng et al.ICLR 2026 · 4 citations
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee et al.ACL 2024 · 20 citations
- Online Preference Alignment for Language Models via Count-based ExplorationChenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang et al.ICLR 2025
- Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference ModelJunshu Pan, Wei Shen, Shulin Huang, Qiji Zhou et al.AAAI 2026 · 7 citations
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyShusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye et al.ICML 2024 · 274 citations
