Lune

NeurIPS2025Top-tier venue

Robust LLM Alignment via Distributionally Robust Direct Preference Optimization

Zaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil, Rahul Jain, Deepak Ramachandran

2025Year
18Citations
6Top-tier citations

Abstract

A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, assuming that they accurately represent real-world user preferences. However, user preferences vary significantly across geographical regions, demographics, linguistic patterns, and evolving cultural trends. This preference distribution shift leads to catastrophic alignment failures in many real-world applications. We address this problem using the principled framework of distributionally robust optimization, and develop two novel distributionally robust direct preference optimization (DPO) algorithms, namely, Wasserstein DPO (WDPO) and Kullback-Leibler DPO (KLDPO). We characterize the sample complexity of learning the optimal policy parameters for WDPO and KLDPO. Moreover, we propose scalable gradient descent-style learning algorithms by developing suitable approximations for the challenging minimax loss functions of WDPO and KLDPO. Our empirical experiments using benchmark data sets and LLMs demonstrate the superior performance of WDPO and KLDPO in substantially improving the alignment when there is a preference distribution shift.

† Work done as postdoctoral researcher at the California Institute of Technology.

39th Conference on Neural Information Processing Systems (NeurIPS 2025). (Zhao et al., 2024;Durmus et al., 2024). Standard preference-learning methods tend to skew toward the preferences represented in the majority of training data, disproportionately penalizing minority opinions and reinforcing biases (Chakraborty et al., 2024). (ii) Reward hacking: The quality of human preference feedback is inherently noisy, ambiguous, and inconsistent, as they are collected from human annotators who may lack domain expertise, exhibit labeling fatigue, or hold conflicting opinions (Zhang et al., 2025;Wu et al., 2025), which can often lead to misaligned preference estimation. This issue is exacerbated by reward hacking, where models learn undesirable shortcuts to maximize the estimated reward function, generating responses that appear aligned but deviate from genuine human intent (Amodei et al., 2016;Skalse et al., 2022;Eisenstein et al., 2024). (iii) Distribution shift: Alignment algorithms use static preference datasets for training, collected under controlled conditions. However, the preferences of real-world users can often be out-of-distribution from that of the training data, depending on the geographical region, demography, linguistic patterns, and emerging social trends, among many others. A model aligned using a specific fixed dataset may fail catastrophically when deployed to users whose preference distribution does not match that of the training data (

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f0ab2b10-0ca8-4614-9a6e-c2096e5881bd

Cited by top-tier papers6

Ask how each one uses it

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines