Perplexity-aware Correction for Robust Alignment with Noisy Preferences
Keyi Kong, Xilie Xu, Di Wang, Jingfeng Zhang, Mohan S. Kankanhalli
Abstract
Alignment techniques are critical in ensuring that large language models (LLMs) output helpful and harmless content by enforcing the LLM-generated content to align with human preferences. However, the existence of noisy preferences (NPs), where the responses are mistakenly labelled as chosen or rejected, could spoil the alignment, thus making the LLMs generate useless and even malicious content. Existing methods mitigate the issue of NPs from the loss perspective by adjusting the alignment loss based on a clean validation dataset. Orthogonal to these loss-oriented methods, we propose perplexity-aware correction (PerpCorrect) from the data perspective for robust alignment which detects and corrects NPs based on the differences between the perplexity of the chosen and rejected responses (dubbed as PPLDiff). Intuitively, a higher PPLDiff indicates a higher probability of the NP because a rejected/chosen response which is mistakenly labelled as chosen/rejected is less preferable to be generated by an aligned LLM, thus having a higher/lower perplexity. PerpCorrect works in three steps: (1) PerpCorrect aligns a surrogate LLM using the clean validation data to make the PPLDiff able to distinguish clean preferences (CPs) and NPs. (2) PerpCorrect further aligns the surrogate LLM by incorporating the reliable clean training data whose PPLDiff is extremely small and reliable noisy training data whose PPLDiff is extremely large after correction to boost the discriminatory power. (3) Detecting and correcting NPs according to the PPLDiff obtained by the aligned surrogate LLM to obtain a
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
- Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMsShangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang et al.ICLR 2026 · 10 citations
- Scalable Valuation of Human Feedback through Provably Robust Model AlignmentMasahiro Fujisawa, Masaki Adachi, Michael A. OsborneNeurIPS 2025 · 6 citations
- Towards Understanding Valuable Preference Data for Large Language Model AlignmentZizhuo Zhang, Qizhou Wang, Shanshan Ye, Jianing Zhu et al.ICLR 2026 · 6 citations
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry ControlYonghui Yang, Wenjian Tao, Jilong Liu, Xingyu Zhu et al.ICML 2026 · 4 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
Related papers
- ROPO: Robust Preference Optimization for Large Language ModelsXize Liang, Chao Chen, Shuang Qiu, Jie Wang et al.ICML 2025
- Unbiased Alignment for Large Language Models with Noisy PreferencesJialiang Wang, Xianming Liu, Xiong Zhou, Hui Liu et al.ICML 2026
- Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisKai Chen, Chunwei Wang, Kuo Yang, Jianhua Han et al.ICLR 2024 · 47 citations
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM AlignmentXiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long et al.ICLR 2026 · 4 citations
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 3 citations
