Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
Guanyu Zhou, Yibo Yan, Xin Zou, Kun Wang, Aiwei Liu, Xuming Hu
摘要
Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination. These biases arise from the visual encoder and the Large Language Model (LLM) backbone, affecting the attention mechanism responsible for aligning multimodal inputs. Existing decoding-based mitigation methods focus on statistical correlations and overlook the causal relationships between attention mechanisms and model output, limiting their effectiveness in addressing these biases. To tackle this issue, we propose a causal inference framework termed CAUSALMM that applies structural causal modeling to MLLMs, treating modality priors as a confounder between attention mechanisms and output. Specifically, by employing back-door adjustment and counterfactual reasoning at both the visual and language attention levels, our method mitigates the negative effects of modality priors and enhances the alignment of MLLM's inputs and outputs, with a maximum score improvement of 65.3% on 6 VLind-Bench indicators and 164 points on MME Benchmark compared to conventional methods. Extensive experiments validate the effectiveness of our approach while being a plug-and-play solution. Our code is available at: https://github.com/The-Martyr/CausalMM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language ModelsHaolang Lu, Yilian Liu, Jingxin Xu, Guoshun Nan 等NeurIPS 2025 · 被引用 23 次
- Hallucination Begins Where Saliency DropsXiaofeng Zhang, Yuanchao Zhu, Chaochen Gu, Xiaosong Yuan 等ICLR 2026 · 被引用 11 次
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsJiaye Qian, Ge Zheng, Yuchen Zhu, Sibei YangNeurIPS 2025 · 被引用 11 次
- Unbiased Missing-Modality Multimodal LearningRuiting Dai, Chenxi Li, Yandong Yan, Lisi Mo 等ICCV 2025 · 被引用 8 次
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper23
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim 等ICLR 2024 · 被引用 354 次
- Deep Structural Causal Models for Tractable Counterfactual InferenceNick Pawlowski, Daniel Coelho de Castro, Ben GlockerNeurIPS 2020 · 被引用 353 次
相关 Paper
- Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMsWei-Yao Wang, Zhao Wang, Helen Suzuki, Yoshiyuki KobayashiICML 2026 · 被引用 8 次
- CausalLens: Sensitivity-Guided Multi-Head Causal Intervention for Hallucination Mitigation in Large Vision-Language ModelsJunyang Ji, Qifan Liu, Wenming Yang, Zhihai HeCVPR 2026
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma 等CVPR 2026 · 被引用 23 次
- Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image GuidanceHaozhe Zhao, Shuzheng Si, Liang Chen, Yichi Zhang 等EMNLP 2025 · 被引用 1 次
- Mitigating Multimodal Hallucinations via Gradient-based Self-ReflectionShan Wang, Maying Shen, Nadine Chang, Chuong Nguyen 等CVPR 2026 · 被引用 2 次
