Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
Yuanchen Wu, Lu Zhang, Hang Yao, Junlong Du, Ke Yan, Shouhong Ding, Yunsheng Wu, Xiaoqiang Li
Abstract
Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on the models' response generation, and overlooking the task question itself. This paper discusses the vulnerability of LVLMs in solving counterfactual presupposition questions (CPQs), where the models are prone to accept the presuppositions of counterfactual objects and produce severe hallucinatory responses. To this end, we introduce "Antidote", a unified, synthetic data-driven posttraining framework for mitigating both types of hallucination above. It leverages synthetic data to incorporate factual priors into questions to achieve self-correction, and decouple the mitigation process into a preference optimization problem. Furthermore, we construct "CP-Bench", a novel benchmark to evaluate LVLMs' ability to correctly handle CPQs and produce factual responses. Applied to the LLaVA series, Antidote can simultaneously enhance performance on CP-Bench by over 50%, POPE by 1.8-3.3%, and CHAIR & SHR by 30-50%, all without relying on external supervision from stronger LVLMs or human feedback and introducing noticeable catastrophic forgetting issues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1822216e-94fd-4160-a7f4-ccc624d6927fCited by top-tier papers5
- Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal SteeringShuliang Liu, Songbo Yang, Dong Fang, Sihang Jia et al.ACL 2026 · 9 citations
- KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value SmoothingSiyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun HeCVPR 2026 · 3 citations
- Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination MitigationTiantian Dang, Chao Bi, Shufan Shen, Jinzhe Liu et al.CVPR 2026 · 2 citations
- Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language ModelsChengsheng Zhang, Chenghao Sun, Xinyan Jiang, Wei Li et al.CVPR 2026 · 2 citations
- Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language ModelsYuxuan Liang, Fan Shi, Rui Zhu, Xu Li et al.CVPR 2026
Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level AlignmentPritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami et al.ICLR 2025
- HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction DataQifan Yu, Juncheng Li, Longhui Wei, Liang Pang et al.CVPR 2024 · 29 citations
- Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMsKejia Zhang, Keda Tao, Jiasheng Tang, Huan WangNeurIPS 2025 · 13 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) ModelsYufang Liu, Tao Ji, Changzhi Sun, Yuanbin Wu et al.EMNLP 2024 · 4 citations
