Lune

ICLR2026Top-tier venue

α\alpha-DPO: Robust Preference Alignment for Diffusion Models via α\alpha Divergence

Yang Li, Songlin Yang, Wei Wang, Xiaoxuan Han, Jing Dong

2026Year

Abstract

Diffusion models have demonstrated remarkable success in high-fidelity image generation, yet aligning them with human preferences remains challenging. Direct Preference Optimization (DPO) offers a promising framework, but its effectiveness is critically hindered by noisy data arising from mislabeled preference pairs and individual preference pairs. We theoretically show that existing DPO objectives are equivalent to minimizing the Forward Kullback-Leibler (KL) divergence, whose mass-covering nature makes it intrinsically sensitive to such noise. To address this limitation, we propose α-DPO, which reformulates preference alignment through the lens of α-divergence. This formulation promotes modeseeking behavior and bounds the influence of outliers, thereby enhancing robustness. Furthermore, we introduce a dynamic scheduling mechanism that adaptively adjusts α according to the observed preference distribution, providing data-aware noise tolerance during training. Extensive experiments on synthetic and realworld datasets validate that α-DPO consistently outperforms existing baselines, achieving superior robustness and preference alignment. The code and project page are available at https://github.com/yangli-lab/Diffusion_ alpha-DPO_ICLR2026/ .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 58740b2d-af2c-4540-80c2-d1a7a32ce671

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines