Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs
Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Dongming Zhang, Zhendong Mao
Abstract
Multi-modal Large Language Models (MLLMs) have shown remarkable generative capabilities across multi-modal tasks, yet remain plagued by hallucinations where generated textual contents are semantically inconsistent with the input images. This work reveals that existing multi-modal preference optimization methods exhibit shortcomings at the preference data decoding stage. Specifically, different response tokens exhibit varying degrees of association with visual content, and consequently, their contributions to reducing hallucinations and generating high-quality responses differ. Nevertheless, most existing methods do not distinguish in their treatment, often handling them uniformly. To address this challenge, we propose a novel preference alignment method: Cross-modal Adaptive Token-rewarded Preference Optimization (Cat-PO). Building upon direct preference optimization, Cat-PO calculates hierarchical visual relevance rewards for each response token at global, local, and semantic levels. It then organically integrates these three rewards to construct a smooth reward mechanism and designs an innovative KL-based customized loss for rewarded tokens, thereby enabling fine-grained correction of hallucinatory outputs. Extensive experiments on various base models and evaluation benchmarks demonstrate that our Cat-PO can significantly reduce hallucinations and align with human preferences to enhance the truthfulness of MLLMs. 1 † Corresponding author. 1 https://github.com/gavinzzx/CatPO Recently, RLHF-V Yu et al. (2024) collects segment-level human preference data and performs dense DPO training. TPO Gu et al. (2024) explores token-level information in DPO for LVLMs. V-DPO Xie et al. (2024) pairs response preferences with image-contrast preferences and employs vision-guided DPO to reinforce visual context learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b15b5cf8-5501-4b8c-84d1-4e0ec391219fBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Detecting and Preventing Hallucinations in Large Vision Language ModelsAnisha Gunjal, Jihan Yin, Erhan BasAAAI 2024 · 312 citations
Related papers
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu et al.EMNLP 2024 · 11 citations
- CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMsJinlan Fu, Shenzhen Huangfu, Hao Fei, Xiaoyu Shen et al.ICLR 2025
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMsYuanshuai Li, Yuping Yan, Junfeng Tang, Zeqi Zheng et al.ICML 2026 · 1 citation
- MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference OptimizationKangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu et al.ICML 2025
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference OptimizationWenqi Liu, Xuemeng Song, Jiaxi Li, Yinwei Wei et al.NeurIPS 2025 · 18 citations
