Looking Beyond Text: Reducing Language Bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Haozhe Zhao, Shuzheng Si, Liang Chen, Yichi Zhang, Maosong Sun, Baobao Chang, Minjia Zhang
摘要
Large vision-language models (LVLMs) have achieved impressive results in vision-language tasks. However, LVLMs suffer from hallucinations caused by language bias, which neglects images while over-relying on text. We identify two reasons for the bias: 1). Different training scales between the LLM pretraining and LVLM alignment stage. 2). The learned inference bias due to short-term dependency of text data. Therefore, we propose LACING, designed to address such bias with MuLtimodal DuAlattention MeChanIsm (MDA) aNd Soft-Image Guidance (SIG). Specifically, MDA adopts a parallel dual-attention mechanism that constructs separate attention for visual and text inputs to enhance integration of visual inputs across model. SIG uses a learnable soft visual prompt during training and inference to replace visual inputs, designed to compel LVLMs to prioritize text inputs during inference. Experiments across different model architectures and scales demonstrate that LACING effectively debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without additional resources. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent TaskShuzheng Si, Haozhe Zhao, Kangyang Luo, Gang Chen 等ACL 2026 · 被引用 14 次
- Analyzing and Mitigating Object Hallucination: A Training Bias PerspectiveYifan Li, Kun Zhou, Xin Zhao, Lei Fang 等AAAI 2026 · 被引用 8 次
- Mitigating Hallucinations in Vision-Language Models through Image-Guided Head SuppressionSreetama Sarkar, Yue Che, Alex Gavin, Peter Anthony Beerel 等EMNLP 2025 · 被引用 1 次
- MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-JudgeSua Lee, Sanghee Park, Jinbae ImACL 2026 · 被引用 1 次
- Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information GainSeulbi Lee, Sangheum HwangICML 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
相关 Paper
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie 等CVPR 2025
- Cross-Modal Attention Calibration for LVLM Hallucination MitigationJiaming Li, Jiacheng Zhang, Zequn Jie, Lin Ma 等CVPR 2026 · 被引用 23 次
- Understanding and Mitigating Hallucination in Large Vision-Language Models via Modular Attribution and InterventionTianyun Yang, Ziniu Li, Juan Cao, Chang XuICLR 2025
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual GuidanceXinrong Chen, Xu Chu, Yingmin Qiu, Hengyuan Zhang 等CVPR 2026 · 被引用 8 次
