PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
Cong Chen, Mingyu Liu, Chenchen Jing, Yizhou Zhou, Fengyun Rao, Hao Chen, Bo Zhang, Chunhua Shen
Abstract
This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To address the challenge, we identify the current lack of a metric that finely measures the quality of the caption at the concept level. We hereby introduce HalFscore, a novel metric built upon the language graph that is designed to evaluate both the accuracy and completeness of dense captions at a granular level. Additionally, we identify the root cause of hallucination as the model's over-reliance on its language prior. To address this, we propose PerturboLLaVA, which reduces the model's reliance on the language prior by incorporating adversarially perturbed text during training. This method enhances the model's focus on visual inputs, effectively reducing hallucinations and producing accurate, image-grounded descriptions without incurring additional computational overhead. PerturboLLaVA significantly improves the fidelity of generated captions, outperforming existing approaches in handling multimodal hallucinations and achieving improved performance across general multimodal benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06c8650f-722e-4e3b-8afe-b299539ff726Cited by top-tier papers21
- Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMsKejia Zhang, Keda Tao, Jiasheng Tang, Huan WangNeurIPS 2025 · 13 citations
- MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language ModelsQiyan Zhao, Xiaofeng Zhang, Yiheng Li, Yun Xing et al.ACM MM 2025 · 11 citations
- On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsRosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng et al.ICML 2026 · 10 citations
- Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination MitigationLexiang Tang, Xianwei Zhuang, Bang Yang, Zhiyuan Hu et al.AAAI 2026 · 8 citations
- MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference OptimizationAshutosh Chaubey, Jiacheng Pang, Mohammad SoleymaniCVPR 2026 · 7 citations
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Multi-Modal Hallucination Control by Visual Information GroundingAlessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary et al.CVPR 2024
- Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and CoverageSaehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi et al.ICML 2025
- PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language ModelsNhat Hoang, Minh Vu, My T. Thai, Manish BhattaraiCVPR 2026 · 1 citation
- DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting LayerJinfeng Wei, Xiaofeng ZhangACM MM 2024 · 25 citations
- VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image CaptionsKazuki Matsuda, Yuiga Wada, Shinnosuke Hirano, Seitaro Otsuki et al.EMNLP 2025
