Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
Xinxin Liu, Ming Li, Zonglin Lyu, Yuzhang Shang, Chen Chen
摘要
Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise-images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi-dimensional preferences into binary labels generates conflicting gradient signals that misguide Diffusion Direct Preference Optimization (DPO). To address this, we propose Semi-DPO, a semi-supervised approach that treats consistent pairs as clean labeled data and conflicting ones as noisy unlabeled data. Our method starts by training on a consensus-filtered clean subset, then uses this model as an implicit classifier to generate pseudo-labels for the noisy set for iterative refinement. Experimental results demonstrate that Semi-DPO achieves state-of-the-art performance and significantly improves alignment with complex human preferences, without requiring additional human annotation or explicit reward models during training. We will release our code and models at: https://github.com/L-CodingSpace/semi-dpo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- -DPO: Robust Preference Alignment for Diffusion Models via DivergenceYang Li, Songlin Yang, Wei Wang, Xiaoxuan Han 等ICLR 2026
- Diffusion Model Alignment Using Direct Preference OptimizationBram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou 等CVPR 2024 · 被引用 89 次
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace 等NeurIPS 2025 · 被引用 36 次
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao 等NeurIPS 2025 · 被引用 2 次
- Aligning Text-to-Image Diffusion Models to Human Preference by ClassificationLongquan Dai, Xiaolu Wei, He Wang, Shaomeng Wang 等NeurIPS 2025
