Robust Multimodal Large Language Models Against Modality Conflict
Zongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang Li
摘要
Despite the impressive capabilities of multimodal large language models (MLLMs) in visionlanguage tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing on the conflicts between model responses and inputs, we study the inherent conflicts in inputs from different modalities that place MLLMs in a dilemma and directly lead to hallucinations. We formally define the modality conflict and construct a dataset named Multimodal Modality Conflict (MMMC) to simulate this phenomenon in vision-language tasks. Three methods based on prompt engineering, supervised finetuning, and reinforcement learning are proposed to alleviate the hallucination caused by modality conflict. Extensive experiments are conducted on the MMMC dataset to analyze the merits and demerits of these methods. Our results show that the reinforcement learning method achieves the best performance in mitigating the hallucination under modality conflict, while the supervised finetuning method shows promising and stable performance. Our work sheds light on the unnoticed modality conflict that leads to hallucinations and provides more insights into the robustness of MLLMs. The code and dataset are available at https://github.com/zmzhang2000/ MMMC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual GuidanceXinrong Chen, Xu Chu, Yingmin Qiu, Hengyuan Zhang 等CVPR 2026 · 被引用 8 次
- FINER: MLLMs Hallucinate under Fine-grained Negative QueriesRui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata 等CVPR 2026 · 被引用 3 次
- GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian UpdatingWeimin Shi, Xiang Li, Kaige Li, Junhao Fang 等AAAI 2026 · 被引用 2 次
- Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal ModelsSharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka 等ICLR 2026
- Modality-Decoupled Online Recursive EditingSiyuan Li, Youyuan Zhang, Fangming Liu, Jing LiICML 2026
它引用的顶会 Paper24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- SHARP: Steering Hallucination in LVLMs via Representation EngineeringJunfei Wu, Yue Ding, Guofan Liu, Tianze Xia 等EMNLP 2025
- Hallucination Augmented Contrastive Learning for Multimodal Large Language ModelChaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen 等CVPR 2024
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
- ZINA: Multimodal Fine-grained Hallucination Detection and EditingYuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham NeubigCVPR 2026 · 被引用 5 次
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma 等CVPR 2026 · 被引用 23 次
