LLM Safety Alignment is Divergence Estimation in Disguise
Rajdeep Haldar, Ziyi Wang, Guang Lin, Yue Xing, Qifan Song
摘要
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance-refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Safety Depth in Large Language Models: A Markov Chain PerspectiveChing-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song ChenNeurIPS 2025 · 被引用 2 次
- Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie TrainingChristian Moya, Alex Semendinger, Guang Lin, Elliott ThornleyICML 2026 · 被引用 1 次
- Graph Unlearning Meets Influence-aware Negative Preference OptimizationQiang Chen, Zhongze Wu, Ang He, Xi Lin 等ACM MM 2025 · 被引用 1 次
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced SafetyGeon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim, Honglak Lee 等ICLR 2026 · 被引用 42 次
- Emergent Misalignment is Easy, Narrow Misalignment is HardAnna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel NandaICLR 2026 · 被引用 25 次
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 等ACL 2025
- A Common Pitfall of Margin-based Language Model Alignment: Gradient EntanglementHui Yuan, Yifan Zeng, Yue Wu, Huazheng Wang 等ICLR 2025
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu 等ICLR 2026
