Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling
Qiyuan Deng, Xuefeng Bai, Kehai Chen, Yaowei Wang, Liqiang Nie, Min Zhang
Abstract
Reinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution shift. Current approaches typically address this issue through online sampling from the target policy, which requires significant computational resources. In this paper, we hypothesize that during off-policy training, while the ranking order of output generated by policy changes, their overall distribution remains relatively stable. This stability allows the conversion of the sampling process from the target policy into a computationally efficient reranking of preference data. Building on this hypothesis, we propose a new framework that leverages the model's intrinsic safety judgment capability to extract reward signals, which are then used to calculate label confidence for preference reordering. Extensive experiments and theoretical analysis demonstrate that the proposed method effectively addresses the distribution shift issue, remarkably enhancing the safety performance while avoiding about 300x computational overheads. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f511f37-b127-49a7-bbc4-bb43ca2da9c7Cited by top-tier papers4
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool UseAradhye Agarwal, Gurdit Singh Siyan, Yash Pandya, Joykirat Singh et al.ICML 2026 · 5 citations
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu et al.AAAI 2026 · 4 citations
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen et al.ACL 2026 · 3 citations
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
Related papers
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced SafetyGeon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim, Honglak Lee et al.ICLR 2026 · 42 citations
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy OptimizationXiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou et al.NeurIPS 2025 · 9 citations
- MPO: Multilingual Safety Alignment via Reward Gap OptimizationWeixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu et al.ACL 2025 · 12 citations
- Reinforcement Learning for Large Language Models via Group Preference Reward ShapingHuaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao et al.EMNLP 2025
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy OptimizationYifan Niu, Han Xiao, Dongyi Liu, Nuo Chen et al.ICLR 2026 · 13 citations
