Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu
摘要
The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integrating visual modalities, thereby unlocking transformative cross-modal applications in a variety of real-world scenarios. Despite their impressive performance, VLMs are prone to significant hallucinations, particularly in the form of cross-modal inconsistencies. Building on the success of Reinforcement Learning from Human Feedback (RLHF) in aligning LLMs, recent advancements have focused on applying direct preference optimization (DPO) on carefully curated datasets to mitigate these issues. Yet, such approaches typically introduce preference signals in a bruteforce manner, neglecting the crucial role of visual information in the alignment process. In this paper, we introduce RE-ALIGN, a novel alignment framework that leverages image retrieval to construct a dual-preference dataset, effectively incorporating both textual and visual preference signals. We further introduce rDPO, an extension of the standard direct preference optimization that incorporates an additional visual preference objective during finetuning. Our experimental results demonstrate that RE-ALIGN not only mitigates hallucinations more effectively than previous methods but also yields significant performance gains in general visual question-answering (VQA) tasks. Moreover, we show that RE-ALIGN maintains robustness and scalability across a wide range of VLM sizes and architectures. This work represents a significant step forward in aligning multimodal LLMs, paving the way for more reliable and effective cross-modal applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation LearningChengxuan Qian, Shuo Xing, Li Li, Yue Zhao 等ICLR 2026 · 被引用 42 次
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao 等NeurIPS 2025 · 被引用 41 次
- SD-MVS: Segmentation-Driven Deformation Multi-View Stereo with Spherical Refinement and EM OptimizationZhenlong Yuan, Jiakai Cao, Zhaoxin Li, Hao Jiang 等AAAI 2024 · 被引用 38 次
- MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View StereoZhenlong Yuan, Cong Liu, Fei Shen, Zhaoxin Li 等AAAI 2025 · 被引用 22 次
- DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View StereoZhenlong Yuan, Jinguo Luo, Fei Shen, Zhaoxin Li 等AAAI 2025 · 被引用 19 次
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 等CVPR 2026 · 被引用 171 次
相关 Paper
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu 等EMNLP 2024 · 被引用 11 次
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMsYuanshuai Li, Yuping Yan, Junfeng Tang, Zeqi Zheng 等ICML 2026 · 被引用 1 次
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMsJinlan Fu, Shenzhen Huangfu, Hao Fei, Xiaoyu Shen 等ICLR 2025
- Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMsZhixiao Zheng, Zheren Fu, Zhiyuan Yao, Dongming Zhang 等ICLR 2026
