MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
Qiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing, Xiaosong Yuan, Feilong Tang, Sinan Fan, Xuhang Chen, Da-Han Wang, Xu-Yao Zhang
摘要
Hallucinations pose a significant challenge in Large Vision Language Models (LVLMs), with misalignment between multimodal features identified as a key contributing factor. This paper reveals the negative impact of the long-term decay in Rotary Position Encoding (RoPE), used for positional modeling in LVLMs, on multimodal alignment. Concretely, under long-term decay, instruction tokens exhibit uneven perception of image tokens located at different positions within the two-dimensional space: prioritizing image tokens from the bottom-right region since in the one-dimensional sequence, these tokens are positionally closer to the instruction tokens. This biased perception leads to insufficient image-instruction interaction and suboptimal multimodal alignment. We refer to this phenomenon as ''image alignment bias.'' To enhance instruction's perception of image tokens at different spatial locations, we propose MCA-LLaVA, based on Manhattan distance, which extends the long-term decay to a two-dimensional, multi-directional spatial decay. MCA-LLaVA integrates the one-dimensional sequence order and two-dimensional spatial position of image tokens for positional modeling, mitigating hallucinations by alleviating image alignment bias. Experimental results of MCA-LLaVA across various hallucination and general benchmarks demonstrate its effectiveness and generality. The code can be accessed in https://github.com/ErikZ719/MCA-LLaVA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language ModelsYang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu 等ACL 2026 · 被引用 9 次
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability DefenseYiyang Huang, Liang Shi, Yitian Zhang, Yi Xu 等ICLR 2026 · 被引用 7 次
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan 等CVPR 2026 · 被引用 3 次
- Inference Time Optimization with Confidence DynamicsYu Wang, Minghao Liu, Jiayun Wang, Jinrui Huang 等ICML 2026 · 被引用 1 次
- Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMsXiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu 等EMNLP 2025
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
相关 Paper
- Mitigating Object Hallucination via Concentric Causal AttentionYun Xing, Yiheng Li, Ivan Laptev, Shijian LuNeurIPS 2024 · 被引用 78 次
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma 等CVPR 2026 · 被引用 23 次
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang 等ICML 2026
- Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level AlignmentPritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami 等ICLR 2025
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang 等EMNLP 2023 · 被引用 344 次
