Debiasing Multimodal Large Language Models via Penalization of Language Priors
Yifan Zhang, Yang Shi, Weichen Yu, Qingsong Wen, Xue Wang, Wenjing Yang, Zhang Zhang, Liang Wang, Rong Jin
摘要
In the realms of computer vision and natural language processing, Multimodal Large Language Models (MLLMs) have become indispensable tools, proficient in generating textual responses based on visual inputs. Despite their advancements, our investigation reveals a noteworthy bias: the generated content is often driven more by the inherent priors of the underlying Large Language Models (LLMs) than by the input image. Empirical experiments underscore the persistence of this bias, as MLLMs often provide confident answers even in the absence of relevant images or given incongruent visual inputs. To rectify these biases and redirect the model's focus toward visual information, we propose two simple, training-free strategies. First, for tasks such as classification or multi-choice question answering, we introduce a ''Post-Hoc Debias'' method using an affine calibration step to adjust the output distribution. This approach ensures uniform answer scores when the image is absent, acting as an effective regularization technique to alleviate the influence of LLM priors. For more intricate open-ended generation tasks, we extend this method to ''Visual Debias Decoding'', which mitigates bias by contrasting token log-probabilities conditioned on a correct image versus a meaningless one. Additionally, our investigation sheds light on the instability of MLLMs across various decoding configurations. Through systematic exploration of different settings, we achieve significant performance improvements-surpassing previously reported results-and raise concerns about the fairness of current evaluation practices. Comprehensive experiments substantiate the effectiveness of our proposed strategies in mitigating biases. These strategies not only prove beneficial in minimizing hallucinations but also contribute to the generation of more helpful and precise illustrations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable SegmentationJian Hu, Jiayi Lin, Junchi Yan, Shaogang GongNeurIPS 2024 · 被引用 39 次
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang 等CVPR 2026 · 被引用 8 次
- From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language ModelsYuying Shang, Xinyi Zeng, Yutao Zhu, Xiao Yang 等ACM MM 2025 · 被引用 5 次
- MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language ModelsSangyun Chung, Se Yeon Kim, Youngchae Chee, Yong Man RoCVPR 2026 · 被引用 5 次
- Mitigating Multimodal Hallucinations via Gradient-based Self-ReflectionShan Wang, Maying Shen, Nadine Chang, Chuong Nguyen 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
相关 Paper
- Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language ModelsCe Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma 等ICLR 2025
- ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language ModelsYeji Park, Deokyeong Lee, Junsuk Choe, Buru ChangAAAI 2025 · 被引用 19 次
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual GuidanceXinrong Chen, Xu Chu, Yingmin Qiu, Hengyuan Zhang 等CVPR 2026 · 被引用 8 次
- VCGD: Visual Clue Guided Decoding with Caption Model for Mitigating Hallucination in Multimodal Large Language ModelsGuoqing Chen, Fu Zhang, Bingqian Liu, Chenglong Lu 等AAAI 2026
- Inter: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance SamplingXin Dong, Shichao Dong, Jin Wang, Jing Huang 等ICCV 2025
