HPS: Hard Preference Sampling for Human Preference Alignment
Xiandong Zou, Wanyu Lin, Yuchen Li, Pan Zhou
摘要
Aligning Large Language Model (LLM) responses with human preferences is vital for building safe and controllable AI systems. While preference optimization methods based on Plackett-Luce (PL) and Bradley-Terry (BT) models have shown promise, they face challenges such as poor handling of harmful content, inefficient use of dispreferred responses, and, specifically for PL, high computational costs. To address these issues, we propose Hard Preference Sampling (HPS), a novel framework for robust and efficient human preference alignment. HPS introduces a training loss that prioritizes the most preferred response while rejecting all dispreferred and harmful ones. It emphasizes "hard" dispreferred responses -those closely resembling preferred ones -to enhance the model's rejection capabilities. By leveraging a single-sample Monte Carlo sampling strategy, HPS reduces computational overhead while maintaining alignment quality. Theoretically, HPS improves sample efficiency over existing PL methods and maximizes the reward margin between preferred and dispreferred responses, ensuring clearer distinctions. Experiments on HH-RLHF and PKU-Safety datasets validate HPS's effectiveness, achieving comparable BLEU and reward scores while greatly improving reward margins and thus reducing harmful content generation. The source code is available at https://github.com/LVLab-SMU/HPS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- Strong Preferences Affect the Robustness of Preference Models and Value AlignmentZiwei Xu, Mohan S. KankanhalliICLR 2025
- Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference OptimizationJiahui Zhu, Yuanjie Shi, Xiyue Peng, Xin Liu 等ICLR 2026
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language ModelsRei Higuchi, Taiji SuzukiICML 2025
- Learning Preference Model for LLMs via Automatic Preference Data GenerationShijia Huang, Jianqiao Zhao, Yanyang Li, Liwei WangEMNLP 2023 · 被引用 3 次
- Preference Ranking Optimization for Human AlignmentFeifan Song, Bowen Yu, Minghao Li, Haiyang Yu 等AAAI 2024 · 被引用 357 次
