LLM Safety Alignment is Divergence Estimation in Disguise
Rajdeep Haldar, Ziyi Wang, Guang Lin, Yue Xing, Qifan Song
Abstract
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance-refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cadc751d-359e-4ce4-8ef7-c5507cc18449Cited by top-tier papers4
- Safety Depth in Large Language Models: A Markov Chain PerspectiveChing-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song ChenNeurIPS 2025 · 2 citations
- Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie TrainingChristian Moya, Alex Semendinger, Guang Lin, Elliott ThornleyICML 2026 · 1 citation
- Graph Unlearning Meets Influence-aware Negative Preference OptimizationQiang Chen, Zhongze Wu, Ang He, Xi Lin et al.ACM MM 2025 · 1 citation
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng et al.ACL 2026 · 1 citation
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
Related papers
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced SafetyGeon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim, Honglak Lee et al.ICLR 2026 · 42 citations
- Emergent Misalignment is Easy, Narrow Misalignment is HardAnna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel NandaICLR 2026 · 25 citations
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen et al.ACL 2025
- A Common Pitfall of Margin-based Language Model Alignment: Gradient EntanglementHui Yuan, Yifan Zeng, Yue Wu, Huazheng Wang et al.ICLR 2025
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu et al.ICLR 2026
