Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
Zhihe Yang, Xufang Luo, Dongqi Han, Yunjian Xu, Dongsheng Li
摘要
Hallucination remains a major challenge for Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) has gained increasing attention as a simple solution to hallucination issues. It directly learns from constructed preference pairs that reflect the severity of hallucinations in responses to the same prompt and image. Nonetheless, different data construction methods in existing works bring notable performance variations. We identify a crucial factor here: outcomes are largely contingent on whether the constructed data aligns on-policy w.r.t the initial (reference) policy of DPO. Theoretical analysis suggests that learning from off-policy data is impeded by the presence of KL-divergence between the updated policy and the reference policy. From the perspective of dataset distribution, we systematically summarize the inherent flaws in existing algorithms that employ DPO to address hallucination issues. To alleviate the problems, we propose On-Policy Alignment (OPA)-DPO framework, which uniquely leverages expert feedback to correct hallucinated responses and aligns both the original and expert-revised responses in an on-policy manner. Notably, with only 4.8k data, OPA-DPO achieves an additional reduction in the hallucination rate of LLaVA-1.5-7B: 13.26% on the AMBER benchmark and 5.39% on the Object-Hal benchmark, compared to the previous SOTA algorithm trained with 16k samples.
- The work was conducted during Zhihe Yang's internship at MSRA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao 等NeurIPS 2025 · 被引用 41 次
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma 等CVPR 2026 · 被引用 23 次
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference OptimizationWenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei 等NeurIPS 2025 · 被引用 18 次
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential RecommendationYu Wang, Yonghui Yang, Le Wu, Yi Zhang 等SIGIR 2026 · 被引用 9 次
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMsYuanshuai Li, Yuping Yan, Junfeng Tang, Zeqi Zheng 等ICML 2026 · 被引用 1 次
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video ModelsHaojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 等ICML 2025
- RLSF-V: Mitigating Hallucinations in MLLMs via Fuzzy Semantic Self-FeedbackChanghao He, ShuhaoYan, Shuxian Li, Xi Peng 等ICML 2026
- Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference OptimizationShuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai 等EMNLP 2025 · 被引用 2 次
- Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference OptimizationJiulong Wu, Zhengliang Shi, Shuaiqiang Wang, Jizhou Huang 等EMNLP 2025 · 被引用 1 次
