Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
Zaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil, Rahul Jain, Deepak Ramachandran
摘要
A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, assuming that they accurately represent real-world user preferences. However, user preferences vary significantly across geographical regions, demographics, linguistic patterns, and evolving cultural trends. This preference distribution shift leads to catastrophic alignment failures in many real-world applications. We address this problem using the principled framework of distributionally robust optimization, and develop two novel distributionally robust direct preference optimization (DPO) algorithms, namely, Wasserstein DPO (WDPO) and Kullback-Leibler DPO (KLDPO). We characterize the sample complexity of learning the optimal policy parameters for WDPO and KLDPO. Moreover, we propose scalable gradient descent-style learning algorithms by developing suitable approximations for the challenging minimax loss functions of WDPO and KLDPO. Our empirical experiments using benchmark data sets and LLMs demonstrate the superior performance of WDPO and KLDPO in substantially improving the alignment when there is a preference distribution shift.
† Work done as postdoctoral researcher at the California Institute of Technology.
39th Conference on Neural Information Processing Systems (NeurIPS 2025). (Zhao et al., 2024;Durmus et al., 2024). Standard preference-learning methods tend to skew toward the preferences represented in the majority of training data, disproportionately penalizing minority opinions and reinforcing biases (Chakraborty et al., 2024). (ii) Reward hacking: The quality of human preference feedback is inherently noisy, ambiguous, and inconsistent, as they are collected from human annotators who may lack domain expertise, exhibit labeling fatigue, or hold conflicting opinions (Zhang et al., 2025;Wu et al., 2025), which can often lead to misaligned preference estimation. This issue is exacerbated by reward hacking, where models learn undesirable shortcuts to maximize the estimated reward function, generating responses that appear aligned but deviate from genuine human intent (Amodei et al., 2016;Skalse et al., 2022;Eisenstein et al., 2024). (iii) Distribution shift: Alignment algorithms use static preference datasets for training, collected under controlled conditions. However, the preferences of real-world users can often be out-of-distribution from that of the training data, depending on the geographical region, demography, linguistic patterns, and emerging social trends, among many others. A model aligned using a specific fixed dataset may fail catastrophically when deployed to users whose preference distribution does not match that of the training data (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- Mitigating Mismatch within Reference-based Preference OptimizationSuqin Yuan, Xingrui Yu, Jiyang Zheng, Lei Feng 等ICLR 2026 · 被引用 4 次
- Revisiting Robustness for LLM Safety Alignment via Selective Geometry ControlYonghui Yang, Wenjian Tao, Jilong Liu, Xingyu Zhu 等ICML 2026 · 被引用 4 次
- Robust Preference Alignment via Directional Neighborhood ConsensusRuochen Mao, Yuling Shi, Xiaodong Gu, Jiaheng WeiICLR 2026 · 被引用 2 次
- Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-JudgeWenbo Zhang, Lijinghua Zhang, Liner Xiang, Hengrui CaiICML 2026 · 被引用 1 次
它引用的顶会 Paper25
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
相关 Paper
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerZhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu 等NeurIPS 2024 · 被引用 119 次
- Geometric-Averaged Preference Optimization for Soft Preference LabelsHiroki Furuta, Kuang-Huei Lee, Shixiang Shane Gu, Yutaka Matsuo 等NeurIPS 2024 · 被引用 24 次
- Leveraging robust optimization for llm alignment under distribution shiftsMingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang 等NeurIPS 2025 · 被引用 2 次
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language ModelsRei Higuchi, Taiji SuzukiICML 2025
- What Matters in Data for DPO?Yu Pan, Zhongze Cai, Huaiyang Zhong, Guanting Chen 等NeurIPS 2025 · 被引用 13 次
